构建真实世界信息整合任务基准,评估大模型跨源推理能力
A Benchmark for Deep Information Synthesis
- 设计多阶段数据采集流程,生成需融合多源信息的任务
- 120个任务覆盖7领域67国,最先进模型最高仅得F1 8.97
- 揭示当前模型在复杂推理与防幻觉上的短板,适合研究者优化智能体
基于大语言模型的智能体在执行涉及工具使用(如网页浏览、代码执行、数据分析)的复杂任务时日益普及。然而,现有评估基准未能充分检验其从多源信息中综合提炼见解并进行深层推理的能力。为此,我们提出DEEPSYNTH,一个新型基准,用于评估智能体在真实、耗时问题上的表现,这些问题需要信息收集、整合与结构化推理以得出洞察。DEEPSYNTH包含120个任务,涵盖7个领域和数据来源,覆盖67个国家。该基准通过多阶段数据采集流程构建,要求标注者获取官方数据源、提出假设、进行人工分析,并设计具有可验证答案的任务。在DEEPSYNTH上评估11个最先进的大语言模型和深度研究智能体,其最高F1分数为8.97,LLM-judge指标得分为17.5,凸显了该基准的挑战性。分析表明,当前智能体在幻觉控制与大规模信息空间推理方面存在明显不足,证实DEEPSYNTH是推动未来研究的关键基准。
原文摘要 · Abstract (English)
Large language model (LLM)-based agents are increasingly used to solve complex tasks involving tool use, such as web browsing, code execution, and data analysis. However, current evaluation benchmarks do not adequately assess their ability to solve real-world tasks that require synthesizing information from multiple sources and inferring insights beyond simple fact retrieval. To address this, we introduce DEEPSYNTH, a novel benchmark designed to evaluate agents on realistic, time-consuming problems that combine information gathering, synthesis, and structured reasoning to produce insights. DEEPSYNTH contains 120 tasks collected across 7 domains and data sources covering 67 countries. DEEPSYNTH is constructed using a multi-stage data collection pipeline that requires annotators to collect official data sources, create hypotheses, perform manual analysis, and design tasks with verifiable answers. When evaluated on DEEPSYNTH, 11 state-of-the-art LLMs and deep research agents achieve a maximum F1 score of 8.97 and 17.5 on the LLM-judge metric, underscoring the difficulty of the benchmark. Our analysis reveals that current agents struggle with hallucinations and reasoning over large information spaces, highlighting DEEPSYNTH as a crucial benchmark for guiding future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。