arXiv:2602.22675cs.CL2026-02被引 3

用并行搜索替代深度推理,提升长程研究代理的效率与泛化能力。

Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization

  • 并行获取证据,减少序列推理步骤,降低上下文开销。
  • 在多个基准上达到新高,如GAIA达75.7%,浏览任务减少70.7%推理步数。
  • 适用于问答与开放研究场景,适合高效智能搜索系统开发者。

当前深度研究代理主要通过增加推理深度来提升性能,但在搜索密集型场景中导致高推理成本与延迟,且跨异构研究环境的泛化能力仍受限。本文提出「搜索更多,思考更少」(SMTL)框架,面向长时程智能体搜索,兼顾效率与泛化。SMTL以并行证据获取取代序列推理,在有限上下文预算下实现高效管理。为支持跨任务类型泛化,我们设计统一数据合成管道,涵盖确定性问答与开放式研究任务,并适配相应评估指标。通过监督微调与强化学习端到端训练,该代理在多个基准上表现优异:BrowseComp(48.6%)、GAIA(75.7%)、Xbench(82.0%)、DeepResearch Bench(45.9%)。相比Mirothinker-v1.0,SMTL在最多100次交互步骤下,将BrowseComp平均推理步数减少70.7%,同时提升准确率。

原文摘要 · Abstract (English)

Recent deep research agents primarily improve performance by scaling reasoning depth, but this leads to high inference cost and latency in search-intensive scenarios. Moreover, generalization across heterogeneous research settings remains challenging. In this work, we propose \emph{Search More, Think Less} (SMTL), a framework for long-horizon agentic search that targets both efficiency and generalization. SMTL replaces sequential reasoning with parallel evidence acquisition, enabling efficient context management under constrained context budgets. To support generalization across task types, we further introduce a unified data synthesis pipeline that constructs search tasks spanning both deterministic question answering and open-ended research scenarios with task appropriate evaluation metrics. We train an end-to-end agent using supervised fine-tuning and reinforcement learning, achieving strong and often state of the art performance across benchmarks including BrowseComp (48.6\%), GAIA (75.7\%), Xbench (82.0\%), and DeepResearch Bench (45.9\%). Compared to Mirothinker-v1.0, SMTL with maximum 100 interaction steps reduces the average number of reasoning steps on BrowseComp by 70.7\%, while improving accuracy.

智能搜索推理效率泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。