← Lab bench
EXP-053Result

Which small local model reads UK paperwork without making numbers up

90.4%

of 426 fields read correctly by a 256M-parameter model on one RTX 4070, with 0.17 invented numbers per page

Delivery note, phone photo: the document photo, and SmolDocling's text with correctly read fields highlighted (What SmolDocling read)Delivery note, phone photo: the document photo, and SmolDocling's text with correctly read fields highlighted (Photo)
◂▸
PhotoWhat SmolDocling read

We generated 30 realistic but fictional UK documents (invoices, delivery notes and inspection forms with handwritten entries), each as a flat scan, a phone photo or a rough photo. Then we ran two small open document models locally and checked every field, and every number they wrote that isn't on the page.

What we tried

  • Generated 30 fictional documents with 426 known fields, 139 of them handwritten, in three capture conditions.
  • Ran each model the way its own model card shows: one call per page, greedy decoding, on a local RTX 4070.
  • Scored every field by exact value, and counted every 3+ digit number in the output that appears nowhere on the page.
  • Looked at every failure by eye before trusting the numbers.

What we measured

MeasureTeleOCR (1.4B)SmolDocling (256M)Note
Fields read correctly69.5%90.4%426 fields on 30 pages
Handwritten fields90.6%93.5%TeleOCR reaches 94.7% on the pages where it didn't loop
Numbers not on the page, per page16.80.17
Pages where output ran away6 of 302 of 30Repeated lines, or layout coordinates instead of text
Median time per page17.0 s48.4 s
Peak GPU memory3.2 GB0.9 GB

What went wrong

  • On one invoice TeleOCR switched into a layout mode and wrote 501 box coordinates instead of the text. Every one of its "invented numbers" on that page is a coordinate.
  • On five other pages TeleOCR repeated a line, such as the subtotal, until it hit the length limit.
  • TeleOCR also drops quantity cells from tables, so its typed-field score is low even on good pages.
  • Our first loader treated TeleOCR as a standard Qwen2.5-VL model. It isn't: its text layers use Qwen3-style attention, so it only loads through its own model code.

What happens next

  • PaddleOCR-VL-1.6 once PyTorch is upgraded, and a loop guard (length cap, repetition check, retry) for TeleOCR.
  • Real, redacted paperwork from a partner, with their permission.
  • Field extraction into a spreadsheet, with a confidence flag on anything a person should check.

Built with

  • SmolDocling-256M-preview CDLA-permissive-2.0
  • TeleOCR Apache-2.0
  • docling-core MIT
  • Hugging Face Transformers Apache-2.0