arXiv:2606.13710cs.AIcs.LG2026-06

让AI研究代理通过三角色协同进化,自主完成复杂科研任务。

Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher

论文配图:Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
图 1 · 摘自论文原文
  • 设计提议者、求解者、评判者三角色协同进化框架
  • 8B模型在长文本研究任务上超越32B静态模型,效率更高
  • 适用于需要持续进化的开放科研场景,如自动文献综述

深度研究与智能体演化是实现通用人工智能在真实世界应用的关键任务。前者使智能体能在开放环境中自主检索与整合信息以解决开放性研究问题,但受限于系统静态的深度研究能力;后者通过环境交互积累经验实现能力进化,但其有效性主要验证于有标准答案的可验证任务,难以适配开放性研究。为此,我们提出混合式开放性三重演化框架(HOTE),利用混合模式强化学习,基于网络级知识推动提议者、求解者与评判者的协同演化,迈向开放任务与环境中的自主进化智能体。在三个长篇深度研究基准上的实验证明,经HOTE训练的8B模型超越了最强的静态开放8-32B模型及现有先进深度研究训练方法,且耗时更低,进一步验证了三模块协同演化的必要性。

原文摘要 · Abstract (English)

Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence. The former enables autonomous retrieval and integration of information in open-ended environments to tackle open-ended research tasks, yet it is constrained by the static parametric deep research capabilities of agent systems. The latter allows agents to autonomously interact with the environment to gain experiences that evolve model capabilities. However, its effectiveness has been widely validated only on verifiable tasks with standard answers, leaving a gap with open-ended research tasks. To bridge these two critical tasks, we propose the Hybrid Open-Ended Tri-Evolution (HOTE) framework, which leverages hybrid-mode reinforcement learning to facilitate the collaborative evolution of a proposer, solver and judge based on web-scale knowledge, moving toward autonomous evolving agents in open-ended tasks and environments. Extensive experiments on three long-form deep research benchmarks demonstrate that the 8B model trained via HOTE surpasses the strongest static open 8-32B models as well as those trained by state-of-the-art deep research training methods with less time overhead, and further verify that the evolution of all three modules in HOTE is indispensable.

AI研究智能体演化协同进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。