arXiv:2505.22501cs.CL2025-05EMNLP被引 45

让AI搜索代理自我进化,无需人工标注就能越搜越准。

EvolveSearch: An Iterative Self-Evolving Search Agent

  • 用自迭代框架融合监督微调与强化学习,实现无监督进化。
  • 在7个复杂问答数据集上平均提升4.7%,超越现有最优方法。
  • 适合想构建自主搜索能力的AI系统开发者使用。

大语言模型(LLM)的快速发展推动了智能体在信息检索方面的能力提升,尤其通过集成搜索引擎和浏览器等工具。然而,当前主流方法在开放网络搜索场景中面临挑战:监督微调受限于开放域数据生成,而强化学习虽收敛快,但数据利用效率低。为此,我们提出EvolveSearch,一种新颖的迭代自演化框架,结合监督微调与强化学习,无需外部人类标注推理数据即可增强智能体的网络搜索能力。在7个多跳问答(MHQA)基准测试上的大量实验表明,EvolveSearch在各轮迭代中持续提升性能,最终在7个基准上平均比当前最优方法提升4.7%,为开放网络搜索领域开启了自演化智能体的新可能。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has transformed the landscape of agentic information seeking capabilities through the integration of tools such as search engines and web browsers. However, current mainstream approaches for enabling LLM web search proficiency face significant challenges: supervised fine-tuning struggles with data production in open-search domains, while RL converges quickly, limiting their data utilization efficiency. To address these issues, we propose EvolveSearch, a novel iterative self-evolution framework that combines SFT and RL to enhance agentic web search capabilities without any external human-annotated reasoning data. Extensive experiments on seven multi-hop question-answering (MHQA) benchmarks demonstrate that EvolveSearch consistently improves performance across iterations, ultimately achieving an average improvement of 4.7\% over the current state-of-the-art across seven benchmarks, opening the door to self-evolution agentic capabilities in open web search domains.

智能体搜索优化自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。