测试大模型能否发现思维链被修改,结果发现它们几乎察觉不了。
Can Reasoning Models Detect Changes to their Chains of Thought?

- 用不同方式篡改思维链,观察模型是否能识别
- 检测准确率很低,多数情况无法发现改动
- 对自身或他人思维链的篡改,识别能力差不多
编辑模型的思维链(CoT)有多种用途,例如用更强模型的推理预填充,或移除可能产生不安全输出的步骤。这些干预的效果可能依赖于模型无法察觉改动,否则模型可能改变行为。本文研究近期推理模型在多种条件下检测思维链篡改的能力:推理过程中及之后,以及使用自身或其它模型的思维链进行预填充。结果表明:(i) 模型的检测准确率极低;(ii) 模型难以判断思维链被如何修改;(iii) 模型识别自身与他人思维链改动的能力基本相当。
原文摘要 · Abstract (English)
There are many reasons one may want to edit a model's chain of thought (CoT) -- e.g., to prefill it with reasoning from a stronger model or to remove steps that may yield unsafe outputs. The success of these interventions plausibly depends on a model's inability to notice them, as the model may alter its behavior if it suspects tampering. In this work, we study whether recent reasoning models are able to detect such interventions on their CoTs under a variety of conditions: both during reasoning and after it, and when prefilled both with their own CoTs and with those of other models. Broadly, we find that (i) models exhibit only very modest detection accuracy; (ii) models struggle to identify *how* their CoT was modified; and (iii) models are about as good at detecting changes to their own CoTs as to those of other models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。