构建企业文档解析新基准,评估AI代理的语义准确性。
ParseBench: A Document Parsing Benchmark for AI Agents

- 基于2000页真实企业文档,涵盖五维能力评测
- 主流方法最高得分84.9%,无模型在所有维度均衡表现
- 适合研究文档理解与AI代理自动化系统开发者
AI代理正在改变文档解析的需求,关键在于语义正确性:解析输出必须保留自主决策所需的结构与意义,包括准确的表格结构、精确的图表数据、语义合理的格式化以及视觉定位。现有基准未能充分反映企业自动化场景,依赖狭窄的文档分布和仅关注文本相似度的评估指标,忽略了代理关键失败。我们提出ParseBench,包含约2000页经人工验证的企业文档,覆盖保险、金融和政府领域,围绕五个能力维度设计:表格、图表、内容忠实度、语义格式化和视觉定位。对14种方法(包括视觉语言模型、专用文档解析器和LlamaParse)的测试显示能力分布碎片化:无一方法在全部维度表现稳定。LlamaParse Agentic以84.9%的综合得分位居第一,揭示当前系统仍存在显著能力缺口。数据集与评估代码已开源。
原文摘要 · Abstract (English)
AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and meaning needed for autonomous decisions, including correct table structure, precise chart data, semantically meaningful formatting, and visual grounding. Existing benchmarks do not fully capture this setting for enterprise automation, relying on narrow document distributions and text-similarity metrics that miss agent-critical failures. We introduce ParseBench, a benchmark of ${\sim}2{,}000$ human-verified pages from enterprise documents spanning insurance, finance, and government, organized around five capability dimensions: tables, charts, content faithfulness, semantic formatting, and visual grounding. Across 14 methods spanning vision-language models, specialized document parsers, and LlamaParse, the benchmark reveals a fragmented capability landscape: no method is consistently strong across all five dimensions. LlamaParse Agentic achieves the highest overall score at 84.9%, and the benchmark highlights the remaining capability gaps across current systems. Dataset and evaluation code are available on https://huggingface.co/datasets/llamaindex/ParseBench and https://github.com/run-llama/ParseBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。