arXiv:2606.01075cs.CL2026-06被引 2

自进化语言模型推理仍有明显泛化差距,难追人工标注效果。

On the Generalization Gap in Self-Evolving Language Model Reasoning

论文配图:On the Generalization Gap in Self-Evolving Language Model Reasoning
图 1 · 摘自论文原文
  • 用自生成监督信号闭环训练,仅依赖基础模型和无标签提示集。
  • 多轮批评修正可显著提升性能,但大模型仍落后于人工标注监督。
  • 适合研究自进化机制局限性或探索低成本模型优化的学者。

近期研究表明,大型语言模型可通过自进化(SE)自我提升,利用模型自身生成的监督信号。本文在严格闭环设置下,考察自进化算法仅能访问无标签提示集与基础模型时,其自生成监督信号与理想监督训练的接近程度。我们在统一离线框架中分析四种典型策略:单轮验证、多轮反馈修订、迭代训练与课程学习。实验基于骑士与说谎者(KK)逻辑推理任务,该任务具有确定解、可控难度和清晰的由易到难泛化测试能力。结果表明,自进化虽持续优于基础模型,但过度投入计算资源后性能趋于饱和,且始终存在显著差距。多轮批评-修订策略结合大模型表现优异,Gemma 12B 几乎逼近人工标注监督效果。在真实世界推理基准上,自进化收益也较为有限。整体揭示了闭环自进化的作用边界,表明当前形式下自生成监督仍不充分。

原文摘要 · Abstract (English)

Recent work suggests that large language models (LLMs) can improve through self-evolution (SE), using supervision signals generated by the model itself. In this work, we ask: under a strict closed-loop setup, where the self-evolution algorithm has access only to an unlabeled prompt set and a base model, how close can internally generated supervision come to oracle-supervised training? We analyze four representative strategies in a unified offline self-evolution framework: single-round verification, multi-turn revision with feedback, iterative training, and curriculum learning. Our primary experiments use Knights and Knaves (KK) logical reasoning tasks, which provide deterministic solutions, controlled difficulty levels, and a clean testbed for easy-to-hard generalization. We first show that self-evolution consistently improves over the base model, but plateaus after excessive training compute is invested, and eventually still leaves a non-trivial gap to oracle supervision. We find that multi-turn critic-revision with large models can reach strong self-evolution performance, with Gemma 12B nearly matching oracle-supervised training. Beyond Knights and Knaves, we also evaluate self-evolution on real-world reasoning benchmarks, where gains are also modest. Overall, our results characterize when closed-loop self-evolution can help and show how internally generated supervision remains insufficient under this minimal formulation.

自进化语言模型泛化差距逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。