Sheaf
Scoreboard

Forms · Level 2 Parse: second passes over Mistral OCR

A second pass that corrects Mistral OCR 4.1's reading of a form page before it is stored. Each row is one model and scope over the same Mistral readings.

Key takeaways

  • An Opus 5.5 check lifts Mistral from 634 to 681 of 707 facts across 15 form pages.
  • GPT-6 Luna adds 12 facts at about 3% of the check's cost, mostly boxes, choices, initials and short handwriting.
  • On dense hand-printed tax forms Luna gains little and fails one page; Qwen3.8 Flash matches its gain on fewer pages, fails four and takes about 164 s a page.

The scoreboard

The numbers load from /api/public/scoreboards/forms-level-2.

How it was measured

Every second pass reads Mistral's reading of the page image and the page itself, and returns only the parts it changed. Each change is checked: a part outside the pass's scope, or a field that is no longer a label and a value, is rejected. Readings are scored as on Level 1, fact by fact against each page's truth, by the same grader. Pages 09 to 15 are held out: they were never used to shape the prompt or the flag rules.

Log

Test us with your documents. Send the bundle that currently ruins someone's afternoon, exactly as it arrived.