让大模型在极少人工干预下稳定自进化,提升数学与推理能力。
Guided Self-Evolving LLMs with Minimal Human Supervision
- 用人类标注样例引导生成问题,结合动态难度课程训练解题模型。
- Qwen3-8B-Base在数学任务上比无监督版本提升3.0分,媲美20倍数据量的模型。
- 有效防止概念漂移,适合追求低人力成本自进化系统的研究者。
AI 自进化被视为通向超智能的路径,即模型能自主获取、优化并内化自身学习经验。然而,无监督自进化系统常因概念漂移、多样性崩溃和错误演化而迅速停滞甚至退化。为实现稳定可控的自进化且最小化人工干预,本文提出 R-Few:一种基于轻量级人类监督的引导式自我对弈挑战者-求解器框架。每轮迭代中,挑战者利用少量人类标注样本引导合成问题生成,求解器则在人类与合成数据上联合训练,并采用在线难度驱动的课程机制。在数学与通用推理基准上,R-Few 实现持续迭代提升。例如,Qwen3-8B-Base 在数学任务上相较 R-Zero 提升 +3.0 分,性能达到 General-Reasoner 水平,后者训练数据量为前者的 20 倍。消融实验验证了接地挑战者训练与课程驱动求解器训练的互补性,进一步分析表明,R-Few 能有效缓解漂移,实现更稳定可控的共进化动态。
原文摘要 · Abstract (English)
AI self-evolution has long been envisioned as a path toward superintelligence, where models autonomously acquire, refine, and internalize knowledge from their own learning experiences. Yet in practice, unguided self-evolving systems often plateau quickly or even degrade as training progresses. These failures arise from issues such as concept drift, diversity collapse, and mis-evolution, as models reinforce their own biases and converge toward low-entropy behaviors. To enable models to self-evolve in a stable and controllable manner while minimizing reliance on human supervision, we introduce R-Few, a guided Self-Play Challenger-Solver framework that incorporates lightweight human oversight through in-context grounding and mixed training. At each iteration, the Challenger samples a small set of human-labeled examples to guide synthetic question generation, while the Solver jointly trains on human and synthetic examples under an online, difficulty-based curriculum. Across math and general reasoning benchmarks, R-Few achieves consistent and iterative improvements. For example, Qwen3-8B-Base improves by +3.0 points over R-Zero on math tasks and achieves performance on par with General-Reasoner, despite the latter being trained on 20 times more human data. Ablation studies confirm the complementary contributions of grounded challenger training and curriculum-based solver training, and further analysis shows that R-Few mitigates drift, yielding more stable and controllable co-evolutionary dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。