arXiv:2510.27072cs.LG2025-10被引 6

探究自对弈如何提升大模型推理能力,揭示其机制与局限。

Towards Understanding Self-play for LLM Reasoning

  • 通过自对弈生成问题并自我解答,实现推理能力提升。
  • 发现自对弈中参数更新稀疏、令牌分布熵降低,与性能正相关。
  • 适合研究大模型自我改进机制的学者参考。

近期基于可验证奖励的强化学习(RLVR)推动了大语言模型(LLM)推理的发展,催生了自对弈后训练方法,即模型通过自主生成和求解问题来提升自身能力。尽管自对弈在领域内与领域外均表现出显著增益,其内在机制仍不清晰。本文以绝对零号推理器(Absolute Zero Reasoner)为视角,对比分析自对弈、RLVR与监督微调(SFT)的训练动态,考察参数更新稀疏性、令牌分布熵变化及替代提议者奖励函数的影响。结合pass@k评估,进一步将这些动态与推理性能关联。结果揭示了自对弈区别于其他后训练策略的机制,指出其固有局限,并为未来通过自对弈改进大模型数学推理指明方向。

原文摘要 · Abstract (English)

Recent advances in large language model (LLM) reasoning, led by reinforcement learning with verifiable rewards (RLVR), have inspired self-play post-training, where models improve by generating and solving their own problems. While self-play has shown strong in-domain and out-of-domain gains, the mechanisms behind these improvements remain poorly understood. In this work, we analyze the training dynamics of self-play through the lens of the Absolute Zero Reasoner, comparing it against RLVR and supervised fine-tuning (SFT). Our study examines parameter update sparsity, entropy dynamics of token distributions, and alternative proposer reward functions. We further connect these dynamics to reasoning performance using pass@k evaluations. Together, our findings clarify how self-play differs from other post-training strategies, highlight its inherent limitations, and point toward future directions for improving LLM math reasoning through self-play.

自对弈推理增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。