---
title: Evals
description: An eval measures how exactly an OCR engine reads a page. The engine reads a published test set, each reading is graded against ground truth a person checked, and the counts go to a public scoreboard.
section: Concepts
order: 7
---

## What an eval counts

| Count | The question | Graded by |
| --- | --- | --- |
| Values read | Does the reading state the value where it belongs? | A model that sees the reading's text, never the page. |
| Values boxed | Does the value's bounding box sit on its own table cell or form field? | A program. |

Values boxed is counted on readings of the page image.

## The parts

| Part | What it is |
| --- | --- |
| Test set | One or two pages from a public source, published with their ground truth. |
| Test | One value on a page, named by its row and column, or by its label. |
| Ground truth | The correct value of each test, and the place it is printed. A person checks every one in [Teach](/docs/filing/teach). |
| Setup | An engine and the input it reads: the page's PDF, or its image at 150 DPI. A scoreboard calls it an option. |
| Reading | What the engine returns for a page: its text, and its bounding boxes if it has any. |
| Grader | The model or program that compares a reading with the ground truth. |

## How a run goes

1. Each setup reads each page of the test set.
2. A model gives every test a verdict on the reading's text: **correct**, **wrong** or **missing**.
3. A program finds where each value's bounding box sits: its own cell or field, its row or line, its block only, a wrong place, or not read.
4. The run reports both counts for each setup, with the engine's time and cost a page.

## What keeps it fair

- **Public pages.** A test set is published: each page with a numbered bounding box on every test, and the answers. Example: [test set A](/test-set-cfpb-disclosures-digital-and-image.html).
- **Checked by a person.** No test counts until a person has checked its answer and its place on the page.
- **No model grades its own reading.** Another model grades it.
- **Errors stay apart.** A timeout or a missing key is an error, never a wrong answer.
- **The grader is checked.** One reading per page is graded twice, and the run says how many verdicts changed.
- **Readings vary.** A run can read each page up to 5 times, and reports each repeat's count.
- **No headroom is said.** When the best setup is right on 95% of the tests or more, the run says that the test set no longer separates quality.

## Where the results go

The [scoreboards](/scoreboards.html). Two are built from eval runs:

| Scoreboard | Count |
| --- | --- |
| [Grounding](/scoreboard-grounding.html) | values boxed |
| [OCR accuracy](/scoreboard-ocr-accuracy.html) | values read, with time and cost a page |

A run updates a scoreboard when one of its counts changes. Its key points are then written from the results by rule, with no model. Each scoreboard names its grader, links its test sets and keeps a log.

## Who runs one

Sheaf runs evals in its own organization, with its API key or from Claude. Any other organization's key is refused.

## Example

A test on a closing statement: **Appraisal fee · Borrower-Paid · At Closing** is `$1,250.00`.

| Setup | Its reading | Values read | Values boxed |
| --- | --- | --- | --- |
| Engine A, page image | `$1,250.00` in the Appraisal fee row, under At Closing, with a box on that cell | correct | own cell |
| Engine B, page image | `$1,250.00` in that row, under Before Closing, with a box on the whole table | wrong | block only |
| Engine C, page image | the row is left out | missing | not read |
