arXiv:2608.15191cs.IR2026-08

让智能研究代理学会自我判断,减少无效搜索步骤。

When Deep Research Agents Stagnate: Enhancing Reasoning with Retrieval-Aware Agent Control

论文配图:When Deep Research Agents Stagnate: Enhancing Reasoning with Retrieval-Aware Agent Control
图 1 · 摘自论文原文
  • 引入无监督信号与检索感知控制器,动态优化搜索策略。
  • 在多个模型上平均减少14次搜索调用,最高提升准确率10%。
  • 适合需要高效推理的复杂任务研究者,尤其关注成本与速度。

本文分析了多种深度研究代理(DRAs)的推理轨迹,发现现有代理常出现推理停滞:多数迭代对最终性能贡献极小,且缺乏对自身进展的感知,难以调整搜索策略或决定终止时机。为解决此问题,我们提出一组无监督信号与检索感知代理控制器(RAAC),辅助代理在研究过程中选择最优动作。RAAC融合信息检索中的新颖性与覆盖度原则,使推理轨迹更高效,显著提升整体表现并减少不必要的迭代,从而降低计算成本与延迟。具体在BrowseComp-Plus数据集上,对多种DRAs加入RAAC后,平均减少14次搜索调用,最佳模型在召回率和准确率上均有提升,准确率最高提升10%(平均3%)。

原文摘要 · Abstract (English)

In this paper, we analyze the reasoning trajectories of a variety of DRAs and show that existing agents often suffer from reasoning stagnation: the majority of iterations contribute little or no improvement to final performance, while agents lack awareness of their trajectories and are therefore ineffective at adapting their search strategies or determining when to terminate. To address this issue, we introduce a set of unsupervised signals and a Retrieval-Aware Agent Controller (RAAC), which assists the agent in selecting optimal actions at each stage of the research process. RAAC incorporates key information retrieval principles, namely search novelty and information coverage, resulting in more effective reasoning trajectories that improve overall performance while reducing unnecessary iterations, and consequently cost and latency. Specifically on BrowseComp-Plus and across a large set of DRAs, adding RAAC reduces the number of search calls by an average of 14, significantly improves the best-performing DRA on recall and accuracy, and achieves an accuracy gain of up to 10% (3% on average).

智能代理推理优化信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。