模型自知错误却仍自信表达,这信号是诊断而非修复工具。
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal

- 通过隐藏状态探测推理正确性,首步准确率达0.79
- 错误推理的自信评分4.55/5,与正确时几乎无异
- 该信号可读取但无法用于纠正错误,适合机制可解释性研究
链式思考提示假设生成的推理反映模型内部计算。我们发现这一假设在特定可测量层面是错误的:模型内部能检测自身推理错误,但对外仍表现出高度自信。对隐藏状态的线性探测器在预测推理正确性上达到0.95 AUROC——甚至在第一个推理步骤就达到0.79——而错误推理的口头信心评分(4.55/5)几乎与正确推理(4.87/5)无异。文本表面分类器仅得0.59,显示20个百分点的差距在输出文本中不可见。该隐藏错误感知现象存在于三个模型族(Qwen、Llama、Phi)、1.5B至72B参数规模及强化学习训练的推理模型(DeepSeek-R1,AUROC 0.852)。自然问题是该信号能否修复错误?四种干预手段——激活引导、探测指导的Best-of-N、自我修正、激活修补——均失败;修补导致输出完全失去连贯性。该信号为诊断性而非因果性:是计算质量的读出,而非可调控的杠杆。这划定了机械可解释性的边界:推理中的错误表征本质上不同于以往可编辑的事实知识表征。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) prompting assumes that generated reasoning reflects a model's internal computation. We show this assumption is wrong in a specific, measurable way: models internally detect their own reasoning errors but outwardly express confidence in them. A linear probe on hidden states predicts trace correctness with 0.95 AUROC -- from the very first reasoning step (0.79) -- while verbalized confidence for wrong traces is 4.55/5, nearly identical to correct ones (4.87/5). A text-surface classifier achieves only 0.59 on the same data, confirming a 0.20-point gap invisible in the generated text. This hidden error awareness holds across three model families (Qwen, Llama, Phi), 1.5B-72B parameters, and RL-trained reasoning models (DeepSeek-R1, 0.852 AUROC). The natural question is whether this signal can fix the errors it detects. It cannot. Four interventions -- activation steering, probe-guided best-of-N, self-correction, and activation patching -- all fail; patching destroys output coherence entirely. The signal is diagnostic, not causal: a readout of computation quality, not a lever to redirect it. This delineates a boundary for mechanistic interpretability: error representations during reasoning are fundamentally different from the factual knowledge representations that prior work has successfully edited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。