← Lab bench
EXP-066Live

A century of handwritten rainfall, read cell by cell

610of 610

cells right wherever two open OCR models, run on one graphics card, wrote the same number on ten handwritten Met Office rainfall sheets from the 1860s to the 1950s

A handwritten 1870s rainfall sheet for Tiverton Cove with nearly every cell shaded green for read right
1/4Tiverton Cove, 1870s: PaddleOCR-VL reads 120 of 130 cells right. Green is right, red wrong, grey missed.
  1. 1 · 1870s
  2. 2 · Two readers
  3. 3 · Self-check
  4. 4 · 1950s

Archives, councils and insurers sit on tables written by hand. We took ten real Met Office ten-year rainfall sheets, one from each decade, and had four open OCR models read every page on one RTX 4070. Each number is checked against the copies made by the volunteers of the Rainfall Rescue project, and a cell only counts when the right number sits under the right year. The live page shows every sheet with each cell coloured right, wrong or missed; hover a cell to see what each model read. Two ideas make the reading usable: the sheet checks itself, because each year's months must add up to the total written at the bottom, and two models that fail on different sheets check each other.

What we tried

  • Real archive pages, not synthetic ones: Met Office ten-year rainfall sheets scanned for Rainfall Rescue, with volunteer transcriptions as the answer key. One sheet per decade, 1860s to 1950s, 1,131 handwritten cells in all.
  • Four open models on one RTX 4070, each with its maker's own prompt, the same page size and the same output limit: PaddleOCR-VL 1.6, WeVisDoc-2B, TeleOCR and SmolDocling.
  • A strict rule: a cell counts only when the right number sits under the right year, to the hundredth of an inch. A figure slipped into the next year's column is wrong.
  • A self-check from the sheet itself: each year's twelve months are added up and compared with the total the model read at the bottom.
  • A second reader: where the two best models wrote the same number, we measured how often that number was right.

What we measured

MeasureResultNote
PaddleOCR-VL 1.6, right number in the right year75.9%859 of 1,131 cells; 49 s a page
WeVisDoc-2B68.5%775 of 1,131; 120 s a page
TeleOCR / SmolDocling23.2% / 0%SmolDocling returned almost nothing on these sheets
Cells where the two best models agree610 of 610 right54% of all cells
Cells at least one of the two reads right90.5%1,024 of 1,131
Self-check flags that held a real misread77 of 78PaddleOCR-VL on 73 sheets; 77 of 111 bad columns caught
PaddleOCR-VL on the wider set83.7%8,000 cells on 73 sheets

What went wrong

  • PaddleOCR-VL looked stuck: its config turns off caching, so each token recomputed the whole page. Turning caching back on took 16 tokens from 34 s to 1.3 s.
  • Some models loop on empty table cells for thousands of tokens, so every model is stopped at the same limit and scored on what it wrote.
  • The first scoring matched numbers in order along a row, which counted a figure shifted into the wrong year as right. The rule is now the right number in the right year, and the scores dropped to match what the page shows.
  • Qwen3-VL-4B overflowed the 12 GB card and slowed to a crawl, even at half the page size.

What happens next

  • Other hand-filled tables with built-in totals: ledgers, timesheets, meter books, survey forms.
  • A reviewer screen that shows only the cells the models disagree on or the self-check flags.

Built with

  • PaddleOCR-VL 1.6 Apache-2.0
  • WeVisDoc-2B Apache-2.0
  • TeleOCR Apache-2.0
  • SmolDocling CDLA-Permissive-2.0
  • Met Office Rainfall Rescue sheets Open Government Licence v3.0