自进化语言模型推理仍有明显泛化差距,难追人工标注效果。
On the Generalization Gap in Self-Evolving Language Model Reasoning

- 用自生成监督信号闭环训练,仅依赖基础模型和无标签提示集。
- 多轮批评修正可显著提升性能,但大模型仍落后于人工标注监督。
- 适合研究自进化机制局限性或探索低成本模型优化的学者。
近期研究表明,大型语言模型可通过自进化(SE)自我提升,利用模型自身生成的监督信号。本文在严格闭环设置下,考察自进化算法仅能访问无标签提示集与基础模型时,其自生成监督信号与理想监督训练的接近程度。我们在统一离线框架中分析四种典型策略:单轮验证、多轮反馈修订、迭代训练与课程学习。实验基于骑士与说谎者(KK)逻辑推理任务,该任务具有确定解、可控难度和清晰的由易到难泛化测试能力。结果表明,自进化虽持续优于基础模型,但过度投入计算资源后性能趋于饱和,且始终存在显著差距。多轮批评-修订策略结合大模型表现优异,Gemma 12B 几乎逼近人工标注监督效果。在真实世界推理基准上,自进化收益也较为有限。整体揭示了闭环自进化的作用边界,表明当前形式下自生成监督仍不充分。
原文摘要 · Abstract (English)
Recent work suggests that large language models (LLMs) can improve through self-evolution (SE), using supervision signals generated by the model itself. In this work, we ask: under a strict closed-loop setup, where the self-evolution algorithm has access only to an unlabeled prompt set and a base model, how close can internally generated supervision come to oracle-supervised training? We analyze four representative strategies in a unified offline self-evolution framework: single-round verification, multi-turn revision with feedback, iterative training, and curriculum learning. Our primary experiments use Knights and Knaves (KK) logical reasoning tasks, which provide deterministic solutions, controlled difficulty levels, and a clean testbed for easy-to-hard generalization. We first show that self-evolution consistently improves over the base model, but plateaus after excessive training compute is invested, and eventually still leaves a non-trivial gap to oracle supervision. We find that multi-turn critic-revision with large models can reach strong self-evolution performance, with Gemma 12B nearly matching oracle-supervised training. Beyond Knights and Knaves, we also evaluate self-evolution on real-world reasoning benchmarks, where gains are also modest. Overall, our results characterize when closed-loop self-evolution can help and show how internally generated supervision remains insufficient under this minimal formulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。