arXiv:2603.02218cs.LGcs.AI2026-03中稿 · ICML被引 14

让大模型持续自我进化,关键在确保每轮生成的数据都有新信息。

Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain

  • 设计三角色协同系统:出题者、解题者、验证者,推动信息迭代
  • 三类机制协同使学习信息量逐轮提升,突破自对弈瓶颈
  • 适合研究自进化系统、大模型持续学习的学者与工程师

大型语言模型使自演化系统成为可能,但多数现有方案实为快速停滞的自对弈。核心问题在于生成数据未带来可学习信息增量。通过自对弈编程任务实验,我们发现可持续自进化需具备逐轮递增可学习信息的自合成数据流。识别出自演化模型的三元角色:出题者(Proposer)生成任务,解题者(Solver)尝试求解,验证者(Verifier)提供训练信号。提出三种系统设计:不对称共演化解决角色能力弱→强→弱的循环;容量增长匹配参数与推理预算以适应信息上升;主动信息搜寻引入外部上下文和新任务源,防止饱和。三者共同构建可度量的系统级路径,将脆弱的自对弈转化为持续自进化。

原文摘要 · Abstract (English)

Large language models (LLMs) make it plausible to build systems that improve through self-evolving loops, but many existing proposals are better understood as self-play and often plateau quickly. A central failure mode is that the loop synthesises more data without increasing learnable information for the next iteration. Through experiments on a self-play coding task, we reveal that sustainable self-evolution requires a self-synthesised data pipeline with learnable information that increases across iterations. We identify triadic roles that self-evolving LLMs play: the Proposer, which generates tasks; the Solver, which attempts solutions; and the Verifier, which provides training signals, and we identify three system designs that jointly target learnable information gain from this triadic roles perspective. Asymmetric co-evolution closes a weak-to-strong-to-weak loop across roles. Capacity growth expands parameter and inference-time budgets to match rising learnable information. Proactive information seeking introduces external context and new task sources that prevent saturation. Together, these modules provide a measurable, system-level path from brittle self-play dynamics to sustained self-evolution.

自进化大模型自对弈信息增益

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。