ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
ExtractBench is a benchmark for schema-guided document extraction, evaluating value accuracy, record completeness, grounding, and cost. It includes 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types. Metrics include order-insensitive value F1, word-level grounding F1, and page-level grounding F1. LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost.
ExtractBench introduces a benchmark for schema-guided extraction in enterprise workflows, where agents follow user-defined schemas to produce outputs with source grounding. The benchmark covers 4,869 pages from 370 documents across 8 domains and 67 document types, with challenge tags. It scores value accuracy (order-insensitive F1), record completeness, grounding (word- and page-level F1), and cost. Commercial VLMs perform well on short documents but truncate records on long ones; coding agents maintain higher accuracy at higher cost. LlamaExtract Agentic Plus achieves top scores on all metrics, matching coding agent accuracy at lower cost.
The benchmark uniquely combines value accuracy, record completeness, grounding, and cost metrics. It uses a scalable curation pipeline with independent-system agreement, synthetic lists, and human verification. The finding that coding agents retain accuracy on long documents but at high cost, while LlamaExtract Agentic Plus achieves similar accuracy more efficiently, suggests a shift toward specialized extraction agents over general VLMs or costly coding approaches.
Enterprise adoption of schema-guided extraction agents may accelerate if specialized models like LlamaExtract Agentic Plus can deliver high accuracy at lower cost. The benchmark's focus on grounding and completeness addresses compliance and audit needs in regulated industries. The performance gap on long documents indicates a market opportunity for solutions that handle complex, multi-page enterprise documents reliably.
ExtractBench provides a standardized way to evaluate document extraction agents, reducing procurement risk for enterprises. The top-performing model, LlamaExtract Agentic Plus, offers a cost-effective alternative to coding agents, potentially lowering the barrier for automating complex document workflows in finance, legal, and healthcare.
Next signals include potential commercial offerings based on LlamaExtract Agentic Plus, further benchmarks expanding to more document types and languages, and integration of such extraction agents into enterprise RAG and automation platforms. Watch for announcements from LlamaIndex or competitors about productizing these capabilities.