96.2%
All reviewed fields matching the golden record
Docparser Labs · n = 77 fields on one document · 2026-07-24
01 MEASUREMENT NOTES
These are early measurements from one human-reviewed document. They are useful for understanding the current system, not a general accuracy promise.
BENCHMARK METHOD
02 METHOD
We compared each returned value with a human-reviewed golden record. A field counts as a match only when its extracted value exactly matches that record for this benchmark run.
Sample: n = 77 fields on one document.
Date: 2026-07-24.
Scope: one template, one document and the current Gemini Flash extraction path.
03 RESULTS
The broader score includes every field in the response under review. The extraction-only score removes two envelope keys the model was not asked to extract.
96.2%
Docparser Labs · n = 77 fields on one document · 2026-07-24
98.7%
Docparser Labs · same sample · 2026-07-24
Excludes two response-envelope keys that the model was not asked to extract.
13.0 s
Docparser Labs single-template benchmark, Gemini Flash · n = 1 document · 2026-07-24
p95 was 14.1 seconds in the same early benchmark run.
04 GOLDEN RECORDS
A reviewer later corrected a due-date value while creating the benchmark record. The golden was wrong before that human correction; the delivered output was not silently rewritten to match it. The corrected due date became the benchmark's ground truth, while the published measurement stays attached to its own run.
05 MODEL COST
The numbers below show a measured model expense for this run. They are our model cost, not your price.
MEASURED RUN / GEMINI FLASH
1,938 input × $0.30 per million input tokens 3,032 output × $2.50 per million output tokens 06 NOT YET PUBLISHED