arXiv:2606.01393cs.CLcs.AI2026-06

构建专家级文档解析新基准,专测复杂长文档的深层理解能力。

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

论文配图:Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
图 1 · 摘自论文原文
  • 基于解析失败样本筛选挑战性文档,覆盖52个专业领域
  • 含4514页标注数据,平均每篇100页,含6.5万条细粒度标注
  • 揭示现有模型在化学式、乐谱等场景严重失效,适合研究者评估真实场景性能

文档解析与识别是视觉语言模型和文档处理系统的核心能力。然而,现有OCR与文档解析基准在覆盖范围和难度上日益不足:多数聚焦常见文档类型或均匀采样页面,现代解析器在此类任务上表现优异,却缺乏对化学式、乐谱、复杂表格及跨页布局等专家级结构的充分标注。我们提出Dr. DocBench,一个面向专家级文档解析的难度感知基准。该基准基于大规模多语言图书语料库构建,涵盖52个BISAC主题领域,通过解析器失败驱动采样选择具有挑战性的文档,目标为多个先进系统均难以处理的案例。包含4,514页经标注的长文档,平均每篇约100页,提供6.5万条高质量页面级与块级标注,涵盖版面、阅读顺序、层级关系及领域特定视觉内容。对基于流水线的解析器与通用视觉语言模型的评估显示,其在现有基准上的强表现无法迁移至本基准。分析揭示各主题、内容类型与结构属性下均存在显著失败,表明Dr. DocBench是诊断与推进文档智能的综合性测试平台。

原文摘要 · Abstract (English)

Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are increasingly limited in coverage and difficulty: many focus on common document genres or uniformly sampled pages where modern parsers already perform strongly, while offering limited annotation for expert-domain structures such as chemical formula, music notation, complex tables, and cross-page layouts. We introduce Dr. DocBench, a difficulty-aware benchmark for expert-level document parsing. Built from a large-scale multilingual book corpus, Dr. DocBench spans 52 BISAC subject domains and selects challenging documents through parser-failure-based sampling, targeting cases where multiple state-of-the-art systems struggle. It contains 4,514 annotated pages from long documents averaging around 100 pages, with 65k high-quality page- and block-level annotations for layout, reading order, hierarchical relations, and domain-specific visual contents. Evaluations of pipeline-based parsers and general-purpose VLMs show that strong performance on existing benchmarks does not transfer to our expert-level document parsing. Our analysis reveals substantial failures across subjects, content types, and structural attributes, highlighting Dr. DocBench as a comprehensive testbed for diagnosing and advancing document intelligence.

文档解析多模态基准测试长文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。