无需外部评价,让大模型自我进化生成创意内容。
G-Zero: Self-Play for Open-Ended Generation from Zero Data

- 用自生成提示的预测变化作为内在奖励,驱动模型自我改进。
- 在无标注数据下实现持续优化,生成质量显著优于基准方法。
- 适合需要长期迭代、无明确评判标准的创造性任务场景。
自演化大模型在可验证领域表现优异,但在开放性任务中依赖外部语言模型评判会带来能力瓶颈和奖励滥用问题。为此,我们提出G-Zero,一种无需验证器的协同进化框架,实现自主持续改进。核心创新是Hint-δ,通过衡量生成模型在无提示与有自生成提示下的响应预测差异,提供内在奖励信号。利用该信号,提案模型通过GRPO训练,持续生成挑战性问题和有效提示以挖掘生成模型盲区;生成模型则通过DPO同步优化,内化提示引导的改进。理论上,我们在理想化标准DPO版本下证明了最优迭代子最优性保证,前提是提案模型能充分探索且伪标签噪声低。通过完全从内部分布动态中提取监督信号,G-Zero突破了外部评判的能力上限,在不可验证领域提供了可扩展、鲁棒的持续自演化路径。
原文摘要 · Abstract (English)
Self-evolving LLMs excel in verifiable domains but struggle in open-ended tasks, where reliance on proxy LLM judges introduces capability bottlenecks and reward hacking. To overcome this, we introduce G-Zero, a verifier-free, co-evolutionary framework for autonomous self-improvement. Our core innovation is Hint-$δ$, an intrinsic reward that quantifies the predictive shift between a Generator model's unassisted response and its response conditioned on a self-generated hint. Using this signal, a Proposer model is trained via GRPO to continuously target the Generator's blind spots by synthesizing challenging queries and informative hints. The Generator is concurrently optimized via DPO to internalize these hint-guided improvements. Theoretically, we prove a best-iterate suboptimality guarantee for an idealized standard-DPO version of G-Zero, provided that the Proposer induces sufficient exploration coverage and the data filteration keeps pseudo-label score noise low. By deriving supervision entirely from internal distributional dynamics, G-Zero bypasses the capability ceilings of external judges, providing a scalable, robust pathway for continuous LLM self-evolution across unverifiable domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。