构建首个真实文档结构化提取统一评测基准,推动技术落地。
READoc: A Unified Benchmark for Realistic Document Structured Extraction
- 将PDF转为语义丰富的Markdown作为真实任务目标
- 基于3576份真实文档构建数据集,覆盖arXiv/GitHub/Zenodo
- 提供标准化、分段与评分一体化评估套件,适配各类模型
文档结构化提取(DSE)旨在从原始文档中提取结构化内容。尽管已有众多DSE系统出现,其统一评估仍不充分,严重制约领域发展。这主要源于现有基准范式碎片化、局部化。为此,我们提出新型基准READoc,将DSE定义为将非结构化PDF转换为语义丰富Markdown的现实任务。READoc数据集源自arXiv、GitHub和Zenodo的3,576份多样化真实文档。同时,我们开发了包含标准化、分段与评分模块的DSE评估套件S³uite,对主流管道工具、专家视觉模型及通用视觉语言模型进行统一评估。首次揭示当前工作与统一、真实DSE目标间的差距。我们期望READoc能推动未来研究,催生更全面、实用的解决方案。
原文摘要 · Abstract (English)
Document Structured Extraction (DSE) aims to extract structured content from raw documents. Despite the emergence of numerous DSE systems, their unified evaluation remains inadequate, significantly hindering the field's advancement. This problem is largely attributed to existing benchmark paradigms, which exhibit fragmented and localized characteristics. To address these limitations and offer a thorough evaluation of DSE systems, we introduce a novel benchmark named READoc, which defines DSE as a realistic task of converting unstructured PDFs into semantically rich Markdown. The READoc dataset is derived from 3,576 diverse and real-world documents from arXiv, GitHub, and Zenodo. In addition, we develop a DSE Evaluation S$^3$uite comprising Standardization, Segmentation and Scoring modules, to conduct a unified evaluation of state-of-the-art DSE approaches. By evaluating a range of pipeline tools, expert visual models, and general VLMs, we identify the gap between current work and the unified, realistic DSE objective for the first time. We aspire that READoc will catalyze future research in DSE, fostering more comprehensive and practical solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。