研究发现大模型难控制思维链,影响安全监控可靠性。
Reasoning Models Struggle to Control their Chains of Thought
- 设计新评测集测试模型是否能按指令控制思维过程
- 大模型仅2.7%情况下能控制思维链,但61.9%能控制输出
- 模型越大越易控制思维链,强化训练反而降低控制能力
思维链(CoT)监控是检测现代推理模型异常行为和理解其动机的有力工具。但如果模型能操控其思维链内容,将削弱监控有效性。为此,我们提出CoT-Control评测套件,包含需在遵守指令下完成任务的题目,如不使用‘染色体’一词解答遗传学问题。结果表明,推理模型的思维链可控性远低于输出可控性:Claude Sonnet 4.5 在思维链上仅2.7%时间可控制,而输出可控性达61.9%。此外,思维链可控性随模型规模增大而提升,但随强化学习训练、推理时计算量增加及题目难度提高而下降。即使给予激励而非直接指令以规避监控,模型仍难以有效控制思维链,尽管被告知被监控时可控性略有上升。通过对抗性提示优化也无法显著提升可控性。当前结果表明思维链可控性尚非主流风险,但其低可控性的机制仍不清楚。鉴于其对监控有效性的重要性,我们建议前沿实验室持续追踪未来模型的思维链可控性。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability -- CoT controllability -- we introduce the CoT-Control evaluation suite, which includes tasks that require models to solve problems while adhering to CoT instructions, e.g., reasoning about a genetics question without using the word 'chromosome'. We show that reasoning models possess significantly lower CoT controllability than output controllability; for instance, Claude Sonnet 4.5 can control its CoT only 2.7% of the time but 61.9% when controlling its final output. We also find that CoT controllability is higher for larger models and decreases with more RL training, test-time compute, and increased problem difficulty. CoT controllability failures extend even to situations in which models are given incentives (as opposed to direct requests) to evade CoT monitors, although models exhibit slightly higher controllability when they are told they are being monitored. Similarly, eliciting controllability by adversarially optimizing prompts does not meaningfully increase controllability. Our results leave us cautiously optimistic that CoT controllability is currently unlikely to be a failure mode of CoT monitorability. However, the mechanism behind low controllability is not well understood. Given its importance for maintaining CoT monitorability, we recommend that frontier labs track CoT controllability in future models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。