arXiv:2608.15064cs.AI2026-08

构建首个长文档结构解析基准,评估目录层级与图文关系恢复能力

LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents

论文配图:LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents
图 1 · 摘自论文原文
  • 设计跨页目录层级与图文关联关系的标注体系
  • 涵盖2582页文档、3937个标题节点、3258个上下文关系
  • 验证结构信息对长文档问答推理有显著提升作用

将视觉文档转化为机器可读表示是文档智能的基础。现有基准集中于页面级元素识别、阅读顺序、公式识别和表格结构。然而,长文档还需恢复文档级结构,包括跨页目录(TOC)层级重建,以及从图表到其标题、注释、来源的一对多类型关联识别。由于这些结构在现有协议中仅部分覆盖或被包含在更广任务中,现有基准无法直接评估两大关键文档级任务:表目录层级恢复与上下文关系恢复。为此,我们提出 extsc{LongDocBench},包含85份真实世界金融报告、教材与学术论文,共2582页,单文档最长105页。提供人工验证标注:3937个标题节点(平均深度3.55,最大深度9),以及2680个图表对象上的3258个上下文关系。进一步评估了这些结构的下游效用与可恢复性。长文档问答实验表明,人工标注的目录层级与上下文关系均能提升推理表现,二者结合带来互补增益。而主流文档解析器虽在页面级表现良好,但在两项任务上仍显不足。为推动进展,我们公开发布 extsc{LongDocBench} 及其评估协议与可复现测试环境。

原文摘要 · Abstract (English)

Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and figures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing benchmarks cannot directly evaluate two key document-level tasks: \emph{Table-of-Contents Hierarchy Recovery} and \emph{Contextual Relationship Recovery}. To benchmark these two tasks, we introduce \textsc{LongDocBench}, comprising 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships annotated across 2,680 table and figure objects. We further evaluate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits. Meanwhile, representative document parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release \textsc{LongDocBench} and its evaluation protocol and reproducible testbed for advancing document-level structure recovery in long documents.

文档理解长文档结构恢复评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。