{"name":"parse-flows","title":"Parse flows","keyPoints":["OCR parses a page in 0.60 to 0.99 s on sheaf.us: 341 pages in three file sets, 12 pages at once.","From the PDF of a digital page, a small model costs $1.30 to $2.19 per 1,000 pages and the OCR engine $5.00. One page a request, it is no faster: Claude Haiku 5.5 takes 7 to 15 s, GPT-6 Luna 17 to 48 s, the OCR engine 9 to 13 s.","Embedded text, read by Poppler with no model, takes 0.05 s and costs nothing, but gives lines, not tables, and nothing on a scan. Read Embedded Text, the parse option built on it, is an ultra cheap fallback, not timed on sheaf.us yet."],"tables":[{"heading":"Speed on sheaf.us: the whole parse of a file set","columns":["Flow","File set","Pages","Read by the OCR engine","Parse","A page"],"rows":[["OCR","Paramount and Warner Bros.: 3 court filings","57","57","34.5 s","0.60 s"],["OCR","SpaceX IPO: 3 filings","48","48","47.3 s","0.99 s"],["OCR","Musk v. Altman: 15 court filings","236","236","210.5 s","0.89 s"]],"note":"Parse: one parse run of the whole file set, from its start to the last page's bounding boxes, which covers render, reading and grounding. OCR: Mistral OCR 4.1 with page annotation, 12 pages at once."},{"heading":"A digital page read from its PDF: correct values, time, tokens and cost","columns":["Reader","Effort","Correct","Time, fee table","Time, two tables","Tokens in","Tokens out","Per 1,000 pages"],"rows":[["Embedded text (Poppler, no model)","none","20 of 20","0.06 s","0.05 s","none","none","$0"],["Claude Haiku 5.5","low","20 of 20","12.6 s","7.0 s","3,076","3,149","$1.88"],["Claude Haiku 5.5","medium","19 of 20","14.6 s","9.7 s","3,076","3,774","$2.19"],["GPT-6 Luna","low","20 of 20","18.9 s","48.4 s","3,664","1,866","$1.30"],["GPT-6 Luna","medium","19 of 20","18.6 s","17.5 s","3,664","2,274","$1.50"],["Mistral OCR 4.1 with page annotation, from the page image","none","20 of 20","9.4 s","12.9 s","billed by the page","billed by the page","$5.00"]],"note":"Two digital pages of test sets A and B, 10 checked values each: a fee table, and two tables side by side. One reading a page, one page a request, from a development computer on Oct 9 2026, not on sheaf.us. Tokens: the mean of the two pages, reasoning included. Medium is each model's default effort. Each 19 is the same value: a row prints $750.00 under two columns, and the model wrote the second under a wrong column. Embedded text is the PDF's text as Poppler gives it, its layout kept by spacing: lines, not tables."},{"heading":"The OCR flow, step by step","columns":["Step","Tool","What it makes","Order"],"rows":[["1. Render every page","Poppler (pdftoppm), 150 DPI","A PNG image of each page, and a thumbnail","16 pages a batch, one batch after another"],["2. Read each page","Mistral OCR 4.1, sent the PNG image","The page's text as Markdown, and its blocks with their bounding boxes","12 pages at once, starting while step 1 still runs"],["3. Find each word","Poppler's text layer (pdftotext); Tesseract only on a page with no text layer","A bounding box for each printed word","After step 1, while step 2 runs; Tesseract one page after another"],["4. Grounding","Matching, no model","A bounding box for each line, table row and table cell of the text","Once a file, after steps 2 and 3"],["5. Classify","Claude Sonnet 5.5","The documents, their names, and where one file holds several","After the parse"]]},{"heading":"The Read Embedded Text flow, step by step","columns":["Step","Tool","What happens to the page","Order"],"rows":[["1. Render every page","Poppler (pdftoppm), 150 DPI","It becomes a PNG image","16 pages a batch, one batch after another"],["2. Read the PDF's embedded text","Poppler (pdftotext)","Every word of the file, with its bounding box, its line and its paragraph","Once a file, while step 1 runs"],["3. List the pictures","Poppler (pdfimages)","The size of each picture on each page","Once a file, while step 1 runs"],["4. Check each page","Five checks, no model","It passes, or it goes to step 2 of the OCR flow","As each batch is rendered"],["5. Write a page that passed","Plain code, no model","Two files: its text as Markdown, and its paragraphs as blocks with their bounding boxes","At once"],["6. Find each word","The words from step 2","A bounding box for each printed word","After step 1"],["7. Grounding","Matching, no model","A bounding box for each line of the text","Once a file, after steps 5 and 6"]],"note":"A page that passes uses no OCR, no model and no network. Its text is never read from the PNG image: the review room shows the image, and two checks look at it. Classify follows, as in the OCR flow."},{"heading":"The five checks of Read Embedded Text","columns":["Check","It sends the page to the OCR engine when","It reads"],"rows":[["1. Scanned","The page has no words, or one picture covers 80% of it","The words and the pictures"],["2. Content outside the text","40 pixels or more of ink that no word explains: a logo, a signature, a drawn tick box","The PNG image and the words' bounding boxes"],["3. Columns","On over 10% of its lines two neighbouring words stand far apart: a table, a form, two columns","The words' bounding boxes"],["4. Hidden text","A word does not show on the page: white text, or text under a bar","The PNG image and the words' bounding boxes"],["5. Garbled text","1% or more of its characters have no meaning","The characters"]],"note":"The checks run in this order and stop at the first one that sends the page to the OCR engine."}],"scoringMethod":"Timed, not scored: one parse run of a whole file set on sheaf.us, from its start to the last page's bounding boxes, which covers render, reading and grounding. A page's time is that total over the set's pages. No grader and no ground truth: a speed needs neither. The table of a digital page read from its PDF is scored: 20 values on two digital pages of the published test sets A and B, each checked by a person, counted when the reading states the value where it belongs. The grader is Claude Opus 5.5, reading the text alone.","log":[{"date":"261009","line":"Added a digital page read from its PDF: embedded text, GPT-6 Luna and Claude Haiku 5.5 at low and medium effort, with correct values, time, tokens and cost."},{"date":"261009","line":"New board. OCR: 0.60 to 0.99 s a page on sheaf.us, on 341 pages. Read Embedded Text is not timed there yet."}],"updatedAt":"2026-10-09T20:03:04.625Z"}