arXiv:2510.14438cs.CL2025-10ACL被引 4

构建网页信息聚合框架,提升智能体复杂推理能力。

WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models

  • 设计主动探索与逻辑命题生成流程,融合12类组合规则
  • 基于10K可验证问答对训练,使模型推理准确率超GPT-4.1
  • 揭示复杂推理才是研究型智能体的真正性能瓶颈

深度研究智能体的核心在于组合推理能力,即能将分散、异构的信息整合为连贯的逻辑洞察。然而当前系统多依赖检索而缺乏深层推理,成功主要取决于简单实体查找而非多步证据聚合。为此,我们提出数据合成管道WebAggregator,推动智能体范式从检索主导转向组合聚合。该方法首先通过主动探索者收集关联知识,再由组合逻辑命题生成器依据12项组合准则构建复杂问题。基于50,000个网站的10,000个可验证问答对,经拒绝采样构建高质量SFT数据集。在该数据上微调后,模型行为发生根本转变,展现出更精准的组合推理与更低的工具冗余。WebAggregator-32B在GAIA、WebWalkerQA和XBench上超越GPT-4.1,媲美Claude-3.7-Sonnet。为弥补现有基准对推理与检索并重的缺失,我们引入WebAggregatorQA测试平台,发现即使具备完美检索能力,顶级模型仍表现不足。结果表明,组合推理才是下一代研究型智能体真正的性能天花板。

原文摘要 · Abstract (English)

The hallmark of Deep Research agents lies in compositional reasoning, the capacity to aggregate distributed, heterogeneous information into coherent logical insights. However, current agentic systems are often retrieval-heavy but reasoning-light, where success is predominantly determined by simple entity-seeking rather than the multi-step aggregation of scattered evidence. To address this, we propose a data synthesis pipeline WebAggregator, designed to shift the agentic paradigm from retrieval-centric to compositional aggregation. Our approach first employs Proactive Explorer to collect interconnected knowledge, then Compositional Logic Proposer to weave knowledge into complex questions using over 12 composition guidelines derived from a rigorous deconstruction of the Deep Research problem setting. By leveraging 10K verifiable QA pairs grounded on 50K websites, we curate a high-quality SFT dataset via rejection sampling. Fine-tuning on this corpus fundamentally transforms agent behavior, fostering deliberate composition reasoning and reduced tool redundancy. The resulting WebAggregator-32B surpasses GPT-4.1 and matches Claude-3.7-Sonnet on GAIA, WebWalkerQA, and XBench. To address the lack of benchmarks that emphasize both reasoning and retrieval, we introduce the WebAggregatorQA testbed, which reveals that even with perfect retrieval, top-tier models still underperformed. These results demonstrate that compositional reasoning, not retrieval, is the true performance ceiling for next-generation research agents.

智能体组合推理数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。