arXiv:2602.18297cs.LGcs.AI2026-02被引 5

用信息论分析推理监控机制,提升大模型推理可监测性

Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory

  • 通过信息论揭示监控有效性的必要条件与误差来源
  • 提出两种训练方法,使监控准确率显著提升且避免推理退化
  • 适合关注大模型可解释性与安全对齐的研究者

链式思维(CoT)监控器是基于大语言模型的系统,用于分析推理轨迹以检测输出是否表现出特定属性(如代码生成中的测试作弊行为)。本文通过信息论分析表明,CoT与输出之间存在非零互信息是监控可行的必要条件,但非充分条件。我们识别出两类实际中可能削弱监控性能的近似误差:信息差距(衡量监控器从CoT中提取可用信息的能力)和诱导误差(衡量监控器对最优监控函数的逼近程度)。进一步证明,可通过针对性训练目标系统性提升监控能力。为此,提出两种互补方法:(a) 基于虚拟标签的方法,直接奖励模型生成能最大化监控准确率的CoT;(b) 更实用的无标签方法,最大化输出与CoT之间的条件互信息。在多种环境中验证,两种方法均显著提升监控准确率,且在对抗监控训练时仍能防止CoT退化,从而缓解任务奖励不完善时的奖励黑客问题。

原文摘要 · Abstract (English)

Chain-of-thought (CoT) monitors are LLM-based systems that analyze reasoning traces to detect when outputs may exhibit attributes of interest, such as test-hacking behavior during code generation. In this paper, we use information-theoretic analysis to show that non-zero mutual information between CoT and output is a necessary but not sufficient condition for CoT monitorability. We identify two sources of approximation error that may undermine the performance of CoT monitors in practice: information gap, which measures the extent to which the monitor can extract the information available in CoT, and elicitation error, which measures the extent to which the monitor approximates the optimal monitoring function. We further demonstrate that CoT monitorability can be systematically improved through targeted training objectives. To this end, we propose two complementary approaches: (a) an oracle-based method that directly rewards the monitored model for producing CoTs that maximize monitor accuracy, and (b) a more practical, label-free approach that maximizes conditional mutual information between outputs and CoTs. Across multiple different environments, we show both methods significantly improve monitor accuracy while preventing CoT degeneration even when training against a monitor, thereby mitigating reward hacking when the task reward is imperfectly specified.

大模型监控信息论推理可解释性对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。