构建真实长文档证据链的评测基准,检验模型跨段落推理能力。
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

- 基于文档内在逻辑挖掘自然分散的证据链,不依赖人工植入。
- 覆盖214篇真实文档,包含481个任务,证据分布符合因果时序关系。
- 适合评估大模型在事故报告、文学分析等高阶推理任务中的表现。
复杂问题的答案常需整合散布在长文档不同位置的证据。例如,事故报告中操作条件、设计缺陷与安全检查缺失可能相隔数十节;小说中角色动机可能仅在远距离场景中显露。现有评测多依赖植入式线索或反向构造的多跳链,其分布与原文不符,难以区分模型是真正推理还是受数据分布干扰。本文提出WILDTRACE,涵盖214篇真实长文档(如技术事故报告、冷门文学作品)的481个任务,所有证据链均源自文档自身的因果、时间与叙事逻辑。基于Pearl因果层级和多跳推理分类,定义七类源内证据结构,采用“源优先”构建流程,从文档结构中挖掘候选链,并通过必要性、答案锚定性、评分一致性、污染抗性及可回答性五阶段验证。随着模型承担高风险分析任务,如何在自然分散证据中实现有效推理,已成为长上下文研究的关键挑战。
原文摘要 · Abstract (English)
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed safety check that jointly explain a disaster may appear dozens of sections apart; in a novel, a character's true motive may surface only through scenes far removed from the moment it becomes relevant. This source-internal evidence integration is central to real-world long-document analysis, yet existing benchmarks largely sidestep it. Needle probes, planted facts, and reverse-engineered multi-hop chains embed evidence that may differ from the host text in distribution, placement, or register, making it unclear whether strong performance reflects genuine source reasoning or distributional artifacts. We introduce WILDTRACE, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic. Drawing on Pearl's causal hierarchy and prior multi-hop reasoning typologies, we define seven source-internal evidence geometries that characterize the distinct relational demands of analytical reading in long documents. A source-first construction pipeline mines candidate trails from document structure before writing questions; each item then undergoes multi-stage validation covering clue necessity, answer groundedness, rubric fidelity, contamination resistance and answerability. As models are increasingly entrusted with real-world high-stakes analytical tasks, this gap between accessing information and reasoning over naturally dispersed evidence emerges as a defining challenge for the next stage of long-context research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。