Find the edges.
Detects exactly where one document ends and the next begins inside a massive, unstructured merged packet.
or click to choose
Multiple PDF, PNG, JPEG & TIFF acceptedSelecting files opens your OS picker. Processing continues in Sheaf. No files are uploaded from this page.
Engineered for the document scale of institutions such as (illustrative examples):
Do not prep. Do not sort. Do not clean up. Just point Sheaf at the raw intake and let the engine handle the reality of operational documents.
What Sheaf does while you aren't looking. A complete forensic pipeline that turns a dumped pile into a structured dossier.
One file arrives holding thirty documents and no table of contents. Sheaf decides where each one starts and ends, and covers every page while doing it.
Your policy says what an application needs. Sheaf binds what actually arrived to what was actually required, and names the difference.
Every correction your reviewers make is a decision Sheaf keeps, and then applies. The queue gets shorter each week, and the rules it runs on are yours rather than an industry average.
Detects exactly where one document ends and the next begins inside a massive, unstructured merged packet.
Recognizes the document context by its content, entirely ignoring useless file names or absent metadata.
Groups the newly separated and named documents into logical, policy-driven dossiers.
Reads the fine print, the tables, and the handwritten notes to pull exactly the fields your operation requires.
Never pretends it is certain. Low-confidence extractions are flagged for an operator's final judgment.
The whole packet goes through text and layout analysis before anything is cut, so scans and photographs count the same as clean digital pages.
Looks for the marks a new document leaves behind: a fresh header, a restarted footer, a "page 1 of 6" counter, a change of form entirely.
Places every cut so the runs tile the packet end to end. No page is left sitting between two documents, and none is claimed by both.
Each separated document is typed from its own content, so a run named scan_0042.pdf still comes out as a title commitment.
Your policy is the list of things that must be satisfied. It varies by product, by program, and by applicant, and Sheaf takes it as given.
Everything that arrived is a candidate, including the documents nobody labelled and the ones filed under the wrong application.
Binds each document to the requirement it actually satisfies, with a confidence attached, and abstains rather than guessing when the evidence is thin.
Reports the requirements nothing covered. The chase list writes itself, and it is the same list your auditor would have built.
Work reaches a person for one of two reasons: the evidence fell short of your threshold, or your policy puts a signature on this class of decision. Everything else is already filed.
Each decision keeps the actor, the pages it covered, and what the engine had proposed before anyone touched it — so the disagreements are visible, not just the outcomes.
When the same situation resolves the same way twice, it stops reaching the queue. Sheaf applies the decision from then on, and cites the two that created it.
What is learned here is scoped to your organization. Nothing your reviewers decide leaves it, and nothing from anyone else’s desk arrives in yours.
The file doesn't need to arrive perfect, but the data leaves that way. Ready for your existing infrastructure.
Extracted data is formatted, normalized, and deposited directly into your existing databases, loan origination software, or systems of record.
No new portals to check. Just accurate data, exactly where you expect it to be.
{
"dossier_id": "req_8849-2A",
"document": {
"type": "1040_schedule_c",
"confidence": 0.99
},
"extracted": {
"borrower_name": "Samuel Hayden",
"adjusted_gross": 142500,
"requires_review": false
}
}
One packet in, a list of real documents out — each with the page range it occupies and the type it was recognized as.
The ranges tile the original file end to end, so nothing you were sent goes missing on the way through.
{
"packet_id": "pkt_8849-2A",
"pages_in": 104,
"documents": [
{ "type": "loan_application", "pages": "1-4" },
{ "type": "pay_stub", "pages": "5-6" }
],
"documents_out": 32,
"pages_orphaned": 0
}
Coverage against your checklist, requirement by requirement, with the document that satisfied each one and the confidence behind it.
What's outstanding comes back as a list you can act on, not a percentage you have to interpret.
{
"checklist_id": "chk_8849-2A",
"satisfied": 6,
"outstanding": [
{ "requirement": "hazard_insurance" },
{ "requirement": "gift_letter" }
],
"complete": false
}
The queue comes back as work, not as a dashboard. What needed a person, what a learned rule already settled, and which decisions it was learned from.
Every automatic resolution points back at the human ones behind it, so an auditor can walk from the outcome to the people without leaving the record.
{
"review_id": "rev_8849-2A",
"queued_for_human": 2,
"settled_by_rule": 11,
"rules": [
{ "id": "prior_year_w2", "learned_from": 2 }
],
"named_signer_required": true
}
Every extracted value carries the file, page and span it came from. An answer you cannot walk back to the page is an answer you cannot defend.
Real work arrives as a scan of a print of a fax. A page with no text layer is the ordinary case here, not the exception that gets routed to a person.
Completeness is the question most systems never answer. Sheaf names the requirement that nothing satisfied, which is usually what is holding the file up.
Every automatic outcome points back at the human decisions behind it, so an auditor can walk from a value to a page without leaving the record.
Every request is checked against the role that made it, records age out on a fixed retention window, and each value stays traceable to its page.