Needs review OCR Processed 19 Aug 2026, 20:32
Terms are those shared with the closest training sample — the reason for the match.
No fields extracted. The type has no field schema, or extraction has not run.
pattern Found by a rule reading the page. The score is a prior for that rule — a VIN whose check digit validates is trusted more than a line of text following a label — reduced by how far the value sat from its label. Measured accuracy by band: python calibrate.py. Currently under-confident: values stated at 60% are right about 86% of the time.
python calibrate.py
model Read from the page image by the vision model. The score is not the model's opinion of itself — it is built from checkable evidence: how many of three independent reads agreed, whether the text it quoted can be found on the page, how the value is rendered (crisp, faint, handwritten), and whether a decoy field that appears on no real document was left empty.
not found Neither a rule nor the model could read it, or the value offered was not printed on the page and was discarded. Blank is deliberate: a missing value is visible to you, a wrong one is not. Values are transcribed from the document, never calculated from other fields.