无需人工标注,让模型自进化出更强的解题与判别能力。
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

- 挑战者生成难题,求解者回应,裁判通过响应逻辑自动评分。
- 在可验证和不可验证任务上分别提升4.2和8.0分,持续迭代不退化。
- 适合研究自演化系统、少样本智能体的开发者参考。
自演化语言模型正成为通往超智能的潜在路径,其优势在于降低对人工监督的依赖。尽管在可验证领域已取得显著进展,但不可验证领域的自演化仍鲜有探索。本文提出一种从零数据出发的裁判协同适应框架(J-Zero),支持可验证与不可验证领域中的统一自进化。挑战者与求解者通过对抗交互共同演化:挑战者生成越来越难的任务,求解者则学习提供更高质量的回答。同时,裁判通过已知顺序的偏好对进行协同适应——依据求解者的回答与其单次回答的分解重组结果的对比关系,而非自身打分。J-Zero在可验证和不可验证任务上分别比基线平均提升4.2分和8.0分,并可在至少十轮迭代中持续优化,而基线在两轮后即开始退化。
原文摘要 · Abstract (English)
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。