Parse flows: how a page is read, step by step
The steps of each parse flow, by OCR and from a PDF's own text, and each flow's speed.
Its key points, its tables, how it was scored and its log load from /api/public/scoreboards/parse-flows.
Test documents
No test set: a speed needs no ground truth. Each flow is timed on the public filings of the sample cases on sheaf.us: court filings and SEC filings.
Glossary
| Term | Definition |
|---|---|
| OCR | Optical character recognition: reading the text, tables and marks of a page from its image or PDF. |
| engine | An OCR product or model that reads a page and returns its text, such as Mistral OCR, Landing AI or Reducto. |
| block | One part of a page as an engine returns it: a paragraph, a table, a heading or a figure. |
| bounding box | The rectangle that marks where something is printed on the page. |
| table cell | One value of a table, where a row and a column meet. |
| form field | A label on a form and the value written for it. |
| grounding | Linking each value of a reading to its bounding box on the page. |
| ground truth | The correct answers for a page, one fact per value, checkbox, choice or signature. Every reading is scored against them. |
| test set | The pages a scoreboard is measured on, with their ground truth. |
| grader | The program or model that compares a reading with the ground truth and counts what is right. |
Sheaf