构建首个需数十步推理的真实多文档问答基准,评估大模型真实信息检索能力。
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
- 通过分解式标注流程,人工生成需数十至数百步推理的自然问题
- 前沿大模型在该基准上最高仅达61.2% F1,主要因召回率低与幻觉
- 适合评估和提升大模型在复杂现实任务中的推理与跨文档理解能力
自动化代理(由大语言模型驱动)正成为信息查询的主流工具。然而,现有大模型代理评估基准很少包含自然、信息性且对人类而言耗时的真实问题。为填补这一空白,我们提出MoNaCo——一个包含1,315个自然且耗时的问题的基准,解决这些问题需数十甚至上百个中间步骤,远超现有任何问答基准。为构建MoNaCo,我们开发了分解式标注流程,以大规模获取并人工解答真实世界的耗时问题。前沿大模型在该基准上最高仅达61.2% F1,主要受限于低召回率和幻觉问题。结果凸显了大模型代理在处理真实世界信息检索任务的复杂性与广度方面的局限性。MoNaCo基准、代码库、提示模板及模型预测结果均已公开。
原文摘要 · Abstract (English)
Automated agents, powered by Large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve -- far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks -- with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。