测试大模型在真实专业PDF中的多模态推理能力,发现顶尖模型仅30.7%正确率。
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
- 用专业人士撰写的100个真实问题-文档对构建评测集,覆盖10个专业领域。
- 最先进模型平均通过率仅30.7%,最低为2%,主要错误源于表格错位、图表误读等模式化缺陷。
- 提供细粒度评分体系和能力标签,适合评估文档AI在医疗、法律等领域的实用表现。
专业领域日常工作中大量信息存在于PDF文件中:保险条款、租赁合同、数据手册、临床指南、施工图纸等。现有文档AI评测通常孤立评估OCR、版面分析、图表理解、表格问答或文档VQA等能力。单个任务高分并不反映模型能否回答专业人士实际提出的复杂问题。GDP_pdf基准直接测量这一能力:包含由十类专业人员撰写的100个问题-文档对,仅当至少两个前沿多模态模型以关键方式失败(如错误回答、遗漏关键证据、虚构主张)时才保留问题。每项配有原子级评分标准,可报告分级评分与严格通过率,并按十一项能力分类打标,涵盖文本提取与定位、表格与图表理解、交叉引用、空间推理及对无依据问题的拒绝回答。在17个前沿模型上测试结果:最佳模型通过率为30.7%,最差为2%。多数错误源自少数重复出现的失效模式:表格错位、图表误读、忽略脚注与排除条款、楼层图符号计数错误、扫描噪声、以及取代早期内容的修订条文。
原文摘要 · Abstract (English)
A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, construction plans. Benchmarks for document AI have generally measured the required capabilities in isolation: OCR, layout analysis, chart reasoning, table QA, document VQA. A high score on any one of them does not necessarily reveal whether a model can answer a realistic question that someone in the field would actually ask about a specific PDF. GDP_pdf is a benchmark built to measure this directly. It consists of question-document pairs authored by working professionals in ten fields, and a candidate question was kept only when at least two frontier multimodal models failed it in a way that mattered: a wrong answer, missed decisive evidence, or a fabricated claim, rather than a superficial difference such as style. Each item comes with a rubric of atomic criteria, so we can report a graded rubric score as well as a strict task-level pass rate, and each item is tagged against a taxonomy of eleven capabilities in three tiers, spanning text extraction and grounding, table and chart comprehension, cross-referencing, spatial reasoning, and abstention on unsupported queries. We report results for seventeen frontier models on the 100-item benchmark: the best model passes only 30.7% of the items and the worst passes 2%. Most errors trace back to a small set of recurring loss patterns: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。