arXiv:2608.00036cs.CLcs.AI2026-08

构建超长文档理解新基准,专测专业场景下的多页证据推理能力。

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding

论文配图:XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
图 1 · 摘自论文原文
  • 基于六领域真实文档,设计跨页、跨文档的复杂问答任务
  • 含1519个问题,72.6%需多页证据,36.6%涉及图表,10.9%跨文档
  • 人工验证+类型化规则,支持精准定位系统失败原因

现实文档任务常要求从数百至数千页的年报、法规、临床指南和技术手册中回答问题,部分问题需跨报告比对。可靠长文档理解是法律合规、医疗、金融和工程等领域应用大模型的前提,决策必须可追溯到具体证据页,错误代价高昂——但现有基准仍主要评估短文本或单页问答。我们提出XL-DocBench,一个完全人工验证的超长文档理解基准,包含来自六个专业领域的1,519个保留问题,上下文最长达2,303页。其中1,103个(72.6%)需多证据页,556个(36.6%)涉及表格、图表或图像,165个(10.9%)需跨文档证据。每个问题配有十二类推理标签、专家标注的证据页、类型化验证规则及答案格式,包括218个无答案情形。通过树状引导合成流程结合内容过滤与194名专家全量验证构建。该基准结合超长专业上下文、页面级证据与类型化规则,填补了以往单页、短多页或纯文本长上下文基准的空白,使未来研究能将系统失败归因于检索、证据使用或规则遵循,而非单一排行榜分数。结果表明当前模型在长上下文、多页证据和结构化推理方面仍表现不佳。

原文摘要 · Abstract (English)

Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing related reports. Reliable long-document understanding is therefore a prerequisite for using LLMs in compliance, clinical, financial, and engineering workflows, where decisions must be traceable to specific evidence pages and the cost of an unsupported answer is high -- yet most existing benchmarks still measure short-context or single-page QA. We introduce XL-DocBench, a fully human-verified benchmark for extra-long document understanding, with 1,519 retained questions from six professional domains and contexts up to 2,303 pages. XL-DocBench goes beyond page-level lookup. 1,103 examples (72.6\%) use multiple evidence pages. The final set also includes 556 questions (36.6\%) that use tables, charts, or figures, and 165 questions (10.9\%) that require evidence from multiple documents. Each question has one of twelve reasoning labels, expert-annotated evidence pages, a typed verification rule, and an answer format, including 218 None-answer cases. We build the benchmark with a tree-guided synthesis pipeline followed by artifact filters and full verification by 194 human experts. By coupling extra-long professional contexts with page-level evidence and typed rules, XL-DocBench fills a gap left by prior single-page, short multi-page, or text-only long-context benchmarks, and lets future work attribute system failures to retrieval, evidence use, or rule following rather than to a single leaderboard score. The results show that current systems still struggle with long contexts, multi-page evidence, and structured reasoning over professional documents.

文档理解多页推理专业问答基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。