无监督强化学习提升数学推理,关键在生成简洁确定性答案
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
- 设计内在奖励机制,强制模型生成简洁确定的答案
- 发现模型逻辑先验决定无监督强化学习成败
- 用流形包裹理论解释成功配置的稳定性
尽管基于结果的强化学习显著提升了大语言模型的数学推理能力,但其对计算成本高昂的真实标签依赖带来了严重的可扩展性瓶颈。无监督强化学习通过内在奖励提供可扩展替代方案,却面临训练动态不透明和灾难性不稳定问题,如策略崩溃与奖励欺骗。本文首先设计并评估一系列显式强化简洁确定性生成的内在奖励;其次,通过测试不同基础模型在内在推理能力谱系上的表现,揭示模型的基础逻辑先验如何决定其成功或失败;最后,提出一种新颖的几何诊断视角,表明成功案例具有流形包裹特性。本研究不仅证明强化简洁确定性输出能有效提升数学推理,更揭示了该方法何时失效,并从几何角度解释其原因。
原文摘要 · Abstract (English)
Although outcome-based reinforcement learning (RL) significantly advances the mathematical reasoning capabilities of Large Language Models (LLMs), its reliance on computationally expensive ground-truth annotations imposes a severe scalability bottleneck. Unsupervised RL guided by intrinsic rewards offers a scalable alternative, yet it suffers from opaque training dynamics and catastrophic instability, such as policy collapse and reward hacking. In this paper, we first design and evaluate a suite of intrinsic rewards that explicitly enforce concise and certain generation. Second, to discover the boundaries of this approach, we test base models across a spectrum of intrinsic reasoning capabilities, revealing how a model's foundational logical prior dictates its success or failure. Finally, to demystify why certain configurations stabilize while others collapse, we introduce a novel geometric diagnostic lens, showing that successful cases are enveloped by manifolds. Ultimately, our work goes beyond merely demonstrating that enforcing concise and certain responses successfully boosts mathematical reasoning; we reveal when this unsupervised approach breaks down and geometrically diagnose why.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。