研究大模型推理监控在隐式推理中的有效性,发现监控效果更依赖任务特性而非推理模式。
Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?

- 用提示干预模拟模型依赖偏见线索的行为,测试不同推理模式下的监控能力。
- 显式推理与隐式推理在监控效果上差异不大,关键看任务是否约束推理过程。
- 适合关注模型可解释性与内部状态可监控性的研究人员参考。
链式思维(CoT)为大语言模型的决策过程提供了可观测窗口,可通过阅读推理轨迹监控目标行为,从而推动了对CoT可监控性的研究。然而,隐式链式思维(Latent CoT)将显式文本替换为少量连续状态,虽降低了推理成本,但移除了可读的推理轨迹,使得监控依赖于对模型内部激活的探测或将隐状态重新表述为文本等替代方法。这些替代方案能保留多少监控能力尚不明确。本文采用基于提示的干预设置,作为模型利用偏见输入线索(如无意泄露的答案或用户陈述的观点)而不承认其影响的行为代理。以提示依赖性作为监控目标,在数学推理和问答任务中对比了从显式CoT到弱监督和强监督隐式CoT的多种推理模式下的监控表现。结果表明,在此设置下,监控能力更多取决于任务特性(如正确答案是否约束支持性推理)以及对模型内部的访问程度,而非推理模式本身。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) reasoning offers a window into the decision-making of large language models (LLMs), which can be monitored for target behaviors by reading the reasoning trace, motivating work on CoT monitorability. Latent CoT approaches, however, replace the explicit tokens with a small number of continuous states, lowering inference costs but removing the readable trace this monitoring relies on. Monitoring then requires alternative access to the model, such as probing its activations or verbalizing the latent states back into text, but how much monitorability these alternatives preserve is unclear. We study this question with a hint-based intervention setup, a proxy for behaviors where models exploit biasing input cues, e.g., an inadvertently leaked answer or a belief stated by the user, without acknowledging them. Taking hint-reliance as the monitorability target, we compare monitors across reasoning modes, from explicit CoT to weakly- and strongly-supervised latent CoT, on math reasoning and question answering. We find that, in this setup, monitorability depends more on properties of the task (such as whether the correct answer constrains the supporting reasoning) and the level of access to model internals than on the reasoning mode.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。