让大模型自主探索世界知识,无需外部奖励即可自我进化。
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

- 用内在奖励机制训练模型主动探索并总结未知环境
- 推理时无需人类干预,性能比原版提升20%
- 小模型14B可超越未辅助的Gemini-2.5-Flash
当前大多数智能体的自我演化依赖人工设定的奖励与规则,一旦失去外部指导即停止进化。本文提出一种基于结果的内在奖励机制,通过衡量智能体自生成的世界知识对下游任务成功率的提升程度,训练其在预执行阶段自发探索与归纳环境。该奖励仅用于训练阶段,使模型学会高效探索与总结。推理时完全不依赖外部奖励或人工指令,能自发完成原生自我演化,以适应未知环境。在Qwen3-30B与Seed-OSS-36B上应用该方法,使WebVoyager和WebWalker任务性能提升20%。最显著的是,一个14B的小模型生成的世界知识使其在未受辅助的情况下超越未经调优的Gemini-2.5-Flash,开创了真正可进化的智能体新范式。
原文摘要 · Abstract (English)
Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human guidance, the evolution stops. In this work, we train agents to possess an intrinsic meta-evolution capability to spontaneously learn about unseen environments prior to task execution. To instill this ability, we design an outcome-based reward mechanism that measures how much an agent's self-generated world knowledge improves its success rate on downstream tasks. This reward signal is used exclusively during the training phase to teach the model how to explore and summarize effectively. At inference time, the agent requires no external rewards or human instructions. It spontaneously performs native self-evolution to adapt to unknown environments using its internal parameters. When applied to Qwen3-30B and Seed-OSS-36B, this shift to native evolution yields a 20% performance increase on WebVoyager and WebWalker. Most strikingly, the generated world knowledge even enables a compact 14B Qwen3 model to outperform the unassisted Gemini-2.5-Flash, establishing a new paradigm for truly evolving agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。