RLVR训练中推理透明度可自发提升,但依赖数据多样性。
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
- 通过验证奖励强化学习,让模型推理过程更易监控。
- 数据多样性和指令遵循数据是提升透明度的关键。
- 推理能力提升不等于透明度提高,适合安全审计研究者。
随着大型推理模型(LRMs)的广泛应用,对其思维链(CoT)轨迹进行安全性审计变得至关重要。近期研究发现,在基于可验证奖励的强化学习(RLVR)初期阶段,监控性——即思维链是否忠实且富有信息地反映内部计算——会自发出现,如同‘免费礼物’。本文通过跨模型族和训练领域的系统评估,使这一现象更加具体化。结果表明,该效应并非普遍成立:监控性提升强烈依赖于数据。特别是,我们证明了数据多样性和指令遵循数据在RLVR训练中的关键作用。进一步显示,监控性与模型能力正交——推理性能提升并不意味着透明度增加。机制分析表明,监控性提升主要源于响应分布变尖锐(熵降低)和对提示的关注度增加,而非更强的因果推理依赖。我们还揭示了监控性动态随训练与评估难度变化的规律。这些发现共同构建了对RLVR下监控性涌现的全面理解,明确了其发生条件与限制。
原文摘要 · Abstract (English)
As Large Reasoning Models (LRMs) are increasingly deployed, auditing their chain-of-thought (CoT) traces for safety becomes critical. Recent work has reported that monitorability--the degree to which CoT faithfully and informatively reflects internal computation--can appear as a "free gift" during the early stages of Reinforcement Learning with Verifiable Rewards (RLVR). We make this observation concrete through a systematic evaluation across model families and training domains. Our results show that this effect is not universal: monitorability improvements are strongly data-dependent. In particular, we demonstrate the critical role of data diversity and instruction-following data during RLVR training. We further show that monitorability is orthogonal to capability--improvements in reasoning performance do not imply increased transparency. Through mechanistic analysis, we attribute monitorability gains primarily to response distribution sharpening (entropy reduction) and increased attention to the prompt, rather than stronger causal reliance on reasoning traces. We also reveal how monitorability dynamics vary with controlled training and evaluation difficulty. Together, these findings provide a holistic view of how monitorability emerges under RLVR, clarifying when gains are likely to occur and when they are not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。