发现强化学习让大模型思维变窄,提出方法防止认知坍缩。
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
- 通过采样偏差与语义耦合机制,揭示策略坍缩的根源。
- 在多个逻辑任务上,新方法使泛化能力提升23%以上。
- 适合关注大模型推理可靠性与鲁棒性的研究者。
基于可验证奖励的强化学习(RLVR)是将大语言模型转化为可靠问题求解器的核心范式,尤其在逻辑密集型领域表现突出。尽管其在实践中取得成功,但尚不清楚它是否激发了新能力,还是仅对已有知识分布进行过度锐化。本文通过形式化过锐化现象——即策略坍缩至有限模式,压制有效替代方案——揭示了这一问题。研究发现,有限批次更新会内在地偏向于采样到的模式,通过语义耦合引发全局性坍缩。为此,提出逆成功优势校准以优先处理困难查询,以及基于记忆网络的分布级校准以多样化采样。实验验证,所提策略能有效提升泛化性能。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) is a central paradigm for turning large language models (LLMs) into reliable problem solvers, especially in logic-heavy domains. Despite its empirical success, it remains unclear whether RLVR elicits novel capabilities or merely sharpens the distribution over existing knowledge. We study this by formalizing over-sharpening, a phenomenon where the policy collapses onto limited modes, suppressing valid alternatives. At a high level, we discover finite-batch updates intrinsically bias learning toward sampled modes, triggering a collapse that propagates globally via semantic coupling. To mitigate this, we propose inverse-success advantage calibration to prioritize difficult queries and distribution-level calibration to diversify sampling via a memory network. Empirical evaluations validate that our strategies can effectively improve generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。