Extraction is the index.
Nothing extra runs. The fields pulled when a document was filed are the same fields a question searches, so a file becomes answerable the moment it is approved.
Where were we invoiced above the rate we signed?
Who is working on our sites with expired insurance?
What renews itself before we get the chance to cancel?
Which customers are still buying at last year's prices?
Do not prep. Do not sort. Do not clean up. Just point Sheaf at the raw intake and let the engine handle the reality of operational documents.
The reading is already done, so the asking is free. Ask a question in plain words and Sheaf answers it across everything you have ever put through it — then names the pages it read to get there.
Nothing extra runs. The fields pulled when a document was filed are the same fields a question searches, so a file becomes answerable the moment it is approved.
Type it the way you would say it to a colleague. There is no syntax to learn, no filter panel to assemble, and nothing to set up before the first answer comes back.
What is missing across four thousand files, where two documents contradict each other, what your own average actually is. None of it is written down anywhere until the files are read together.
An answer arrives with the documents behind it and the pages they sat on. Open one and you are looking at the scan itself, not a summary of it.
The file doesn't need to arrive perfect, but the data leaves that way. Ready for your existing infrastructure.
A question comes back as rows, not prose. The same shape every time, ready to drop straight into a spreadsheet, a report, or whatever already runs your week.
Every row carries the documents it was drawn from, so a number in a board pack can be walked back to a page in a scan without leaving the record.
{
"question": "invoiced above signed rate",
"files_read": 3914,
"elapsed_ms": 61,
"rows": [
{
"subject": "Meridian Freight",
"delta_usd": 4120,
"cites": ["RC-4471 p1", "INV-88213 p2"]
}
],
"total_usd": 16765
}