构建真实职场文档推理基准,测试智能体在海量混乱文件中精准定位答案的能力。
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

- 用智能体流水线生成跨文档问题,模拟真实工作场景中的信息搜寻。
- 8个领域共9664份真实文档,372M token总量远超模型上下文窗口。
- 最强模型仅59.4%准确率,证明当前智能体处理复杂档案仍存巨大挑战。
大语言模型正越来越多地作为智能体,在文档中推理而非依赖参数化知识。本文研究档案基础推理:从大量杂乱的职场文件中定位稀疏证据,调和术语、单位和时间格式不一致,并计算答案。现有基准仅覆盖部分场景,且未同时强调档案基础性、智能体探索性和跨领域覆盖。我们提出Agora,包含362个问题与八个领域的9,664份真实文档(共372M tokens),远超任何模型的上下文窗口,迫使智能体必须有策略地探索而非盲目扫描。Agora通过智能体流水线构建,包括跨文档任务合成、防泄漏混淆和难度筛选。评估八个模型发现,该任务远未解决:即使最强模型也仅达59.4%准确率,且各领域表现差异显著。
原文摘要 · Abstract (English)
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating sparse evidence across a large, messy collection of workplace files, reconciling inconsistent terminology, units, and time conventions, and computing an answer. Existing benchmarks address only parts of this setting and none jointly stresses archive-groundedness, agentic exploration, and cross-domain coverage. We introduce Agora, a benchmark pairing 362 questions with eight domain collections of 9,664 authentic documents and 372M tokens, far exceeding any model's context window, so agents must explore deliberately rather than scan exhaustively. Agora is built by an agentic pipeline combining cross-document task synthesis, leakage-preventing obfuscation, and difficulty filtering. Evaluating eight models, we find the task far from solved: even the strongest reaches only 59.4% accuracy, with notable variation across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。