arXiv:2510.27338cs.LG2025-10NeurIPS被引 8

强化学习训练的推理模型会生成难以阅读的思考链,但答案仍正确。

Reasoning Models Sometimes Output Illegible Chains of Thought

  • 用结果驱动的强化学习让模型产生难读的思考过程
  • 强制使用可读部分后准确率下降53%,说明难读链也有效
  • 适合关注模型可解释性与安全监控的研究者

通过基于结果的强化学习(RL)训练语言模型进行链式思考(CoT)推理,已展现出优异性能。监控此类模型的思考链有助于理解其意图并检测潜在恶意行为。然而,有效性依赖于思考链的可读性和真实性。我们研究了14个推理模型的可读性,发现强化学习常导致人类和AI监控器无法理解推理过程,除Claude外,各模型生成的思考链均难以阅读,但最终答案仍完全可读。实验显示,模型利用不可读的推理路径获得正确答案(强制使用可读部分后准确率下降53%),但在重采样时可读性与性能无显著相关性,表明关系更复杂。此外,难题上的可读性下降更明显。我们提出可能原因包括隐写术、训练痕迹及残留标记。结果表明,若不显式优化可读性,基于结果的强化学习会自然产生越来越不透明的推理过程,可能削弱监控手段的有效性。

原文摘要 · Abstract (English)

Language models trained via outcome-based reinforcement learning (RL) to reason using chain-of-thought (CoT) have shown remarkable performance. Monitoring such a model's CoT may allow us to understand its intentions and detect potential malicious behavior. However, to be effective, this requires that CoTs are legible and faithful. We study CoT legibility across 14 reasoning models, finding that RL often causes reasoning to become illegible to both humans and AI monitors, with reasoning models (except Claude) generating illegible CoTs while returning to perfectly readable final answers. We show that models use illegible reasoning to reach correct answers (accuracy dropping by 53\% when forced to use only legible portions), yet find no correlation between legibility and performance when resampling - suggesting the relationship is more nuanced. We also find that legibility degrades on harder questions. We discuss potential hypotheses for these results, including steganography, training artifacts, and vestigial tokens. These results suggest that without explicit optimization for legibility, outcome-based RL naturally produces models with increasingly opaque reasoning processes, potentially undermining monitoring approaches.

链式思考可解释性强化学习模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。