构建多页文档解析基准,覆盖中英文真实场景。
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

- 设计跨页文档端到端评估协议,支持文本/表格/公式等多模态解析。
- 包含433份人工标注文档共3246页,覆盖15类真实文档类型。
- 适合研究文档结构恢复与视觉内容保留的学者与工程师使用。
文档解析将视觉丰富的文档转换为机器可读的结构化表示,是信息系统的重要基础。尽管已有诸多基准,但现有方法仍无法满足真实场景需求:或局限于特定任务,或仅评估单页文本中心设置,且缺乏对语义连续性、层级结构恢复和视觉内容保留的细粒度评估。为此,我们提出MPDocBench-Parse,一个面向真实应用的多页文档解析基准。该基准包含433份人工标注文档,共3246页,涵盖15种中英文文档类型,布局多样,支持文档级端到端评估。我们设计了全面的评估协议,覆盖文本、表格、公式识别,截断文本与表格合并,图示提取,阅读顺序及标题层级恢复。实验表明,现有模型虽在基本文本提取上表现良好,但在语义连续性整合、视觉内容解析和层级结构恢复方面仍存在明显不足。MPDocBench-Parse为推动文档解析向更真实场景演进提供了统一基础。
原文摘要 · Abstract (English)
Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks have been proposed for document parsing, they remain inadequate for realistic scenarios. Existing benchmarks either focus on specific tasks or assess only single-page, text-centric settings, making them insufficient for practical multi-page parsing. Moreover, they lack fine-grained evaluation of semantic continuity, hierarchical structure recovery, and visual content preservation. To address these gaps, we propose MPDocBench-Parse, a benchmark for multi-page document parsing in real-world applications. It contains 433 manually annotated documents with 3,246 pages, covering 15 document types in English and Chinese, with diverse layout styles, and supports document-level end-to-end evaluation. We further design a comprehensive protocol for content fidelity and logical structure, covering text, table, and formula recognition, truncated text and table merging, figure extraction, reading order, and heading hierarchy recovery. Experiments show that, while existing models perform well on basic text extraction, they still suffer clear limitations in semantic continuity integration, visual content parsing, and hierarchical structure recovery. MPDocBench-Parse provides a unified foundation for advancing document parsing toward more realistic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。