arXiv:2412.07626cs.CVcs.AI2024-12CVPR被引 160

构建多类型文档解析新基准,提升评估全面性与公平性

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

论文配图:OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
图 1 · 摘自论文原文
  • 覆盖九类文档来源,含手写笔记、密集排版报纸等复杂场景
  • 支持端到端及细粒度属性分析,涵盖19种版式与15类标签
  • 适用于模型对比与性能诊断,适合文档理解研究者使用

文档内容提取是计算机视觉中的关键任务,支撑大语言模型(LLMs)和检索增强生成(RAG)系统的数据需求。尽管近期取得进展,现有文档解析方法仍因文档类型覆盖有限、评估方式过于简化而缺乏公平与全面的评测。为此,我们提出OmniDocBench,一个包含九类文档源的新型基准,涵盖学术论文、教科书,以及手写笔记、密集排版报纸等高难度场景。该基准支持灵活的多层级评估,从端到端测试到基于任务与属性的细粒度分析,涵盖19种布局类别和15种属性标签。我们对基于流水线的方法与端到端视觉-语言模型进行了全面评估,揭示了其在不同文档类型中的优劣势。OmniDocBench为文档解析的公平、多样与精细化评估设立了新标准。数据集与代码已公开于https://github.com/opendatalab/OmniDocBench。

原文摘要 · Abstract (English)

Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have not been fairly and comprehensively evaluated due to the narrow coverage of document types and the simplified, unrealistic evaluation procedures in existing benchmarks. To address these gaps, we introduce OmniDocBench, a novel benchmark featuring high-quality annotations across nine document sources, including academic papers, textbooks, and more challenging cases such as handwritten notes and densely typeset newspapers. OmniDocBench supports flexible, multi-level evaluations--ranging from an end-to-end assessment to the task-specific and attribute--based analysis using 19 layout categories and 15 attribute labels. We conduct a thorough evaluation of both pipeline-based methods and end-to-end vision-language models, revealing their strengths and weaknesses across different document types. OmniDocBench sets a new standard for the fair, diverse, and fine-grained evaluation in document parsing. Dataset and code are available at https://github.com/opendatalab/OmniDocBench.

文档解析评估基准多类型文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。