Forms · Level 1: OCR engines on form pages
Twenty engine setups reading the same eight form pages, each from the PDF and from the page image, scored on whether every value, box, choice and signature reached the Markdown exactly.
Key takeaways
- Mistral OCR 4.1 with an Opus 5.5 check leads at 385 of 386 facts; Opus 5.5 alone reads 384.
- Handwriting separates the engines: the four digital URLA and IRS pages hardly separate anyone.
- On seven held-out pages, including hand-printed tax forms, Mistral alone reads 265 of 321 facts and the Opus check 297.
The scoreboard
The numbers load from /api/public/scoreboards/forms-level-1.
How it was measured
Each page has a truth file: every fact a correct reading contains, such as a value, a ticked box, a chosen option or a signature, read off the page at full zoom. A grader, Claude Opus 5.5, reads only each setup's text and marks every fact correct, wrong or missing. Image input is the 150 DPI page render Sheaf sends; PDF input is the original page. The pages are sanitized loan documents and NIST Special Database 6 tax forms in synthetic hand print.
Sheaf