ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
ExtractBench is a benchmark for schema‑guided extraction from enterprise documents, covering 4,869 pages across 370 documents, 8 domains and 67 types. It evaluates value accuracy, record completeness, grounding and extraction cost.
- ExtractBench offers the first large‑scale benchmark for schema‑guided enterprise document extraction.
- It evaluates accuracy, completeness, grounding evidence, and extraction cost in a single suite.
- The dataset includes 4,869 pages from 370 documents across eight domains and 67 types.
- An open‑source evaluation toolkit is provided to facilitate reproducible research.
Researchers introduced ExtractBench to address the lack of standardized evaluation for schema‑guided document extraction in enterprise settings. The benchmark comprises 4,869 pages drawn from 370 real‑world documents spanning eight business domains and 67 document types.
ExtractBench uniquely scores four dimensions: value accuracy, record completeness at scale, grounding metadata, and the computational cost of extraction. This multi‑metric approach enables developers to compare models not only on correctness but also on efficiency and traceability.
The benchmark is released alongside an open‑source evaluation framework, allowing the AI community to reproduce results and benchmark new models. Its comprehensive coverage aims to accelerate progress in enterprise‑focused NLP applications.
By providing a shared yardstick, ExtractBench helps organizations assess the readiness of extraction agents for production workflows, reducing the risk of costly deployment failures.
Provides a concrete testbed to build and fine‑tune extraction models with clear performance metrics.
Helps assess the reliability and cost‑effectiveness of AI agents before deployment in critical workflows.
Offers a rich, real‑world dataset for academic projects and thesis work on document understanding.
Shows how AI can automate complex data extraction tasks across industries.
- schema‑guided extraction
- Extracting information from a document according to a user‑defined schema that specifies the desired fields and structure.
Alibaba unveils its most capable AI model to date, not far behind Moonshot’s in size - WTVB
At Colleges, the AI Boom Means Everyone Wants to Dabble in Computer Science - U.S. News & World Report
Education Notebook: Trine University team to tackle artificial intelligence issues through seven-month program - The Journal Gazette
AI reveals a massive algae boom across the world’s oceans - ScienceDaily
EHR-based AI beckons rapid-response team to head off avoidable in-hospital deaths - HealthExec
SecurityDisrupting a Criminal Scam Operation
OpenAI shut down accounts linked to a Cambodia-based criminal network using ChatGPT for romance and investment scams.
INTERPOL report finds AI linked to more than half of cybercrime in Africa - Interpol
A recent INTERPOL report found that AI is linked to more than half of cybercrime cases in Africa.

EU AI Act Article 50: What the 2026 Transparency Rules Mean for AI Teams
The EU AI Act’s Article 50 introduces enforceable transparency rules starting August 2, 2026, requiring AI teams to document and disclose key system details.
Janesville becomes an AI data center battleground - PBS Wisconsin
Janesville is becoming a key location for AI data centers, with major companies competing for space. This development is expected to bring significant investment and job creation to the area.
Potential US ban on Chinese AI models could cost businesses US$12 billion a year - South China Morning Post
A potential US ban on Chinese AI models could cost businesses up to $12 billion per year, according to a report from the South China Morning Post.
Tech: Casar wants to ban AI superintelligence - Punchbowl News
A U.S. representative has introduced a bill to prohibit the development of AI systems smarter than humans, citing existential risks.