医学大模型推理错误可被线性解码,但固定线性修正无效。
Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes

- 通过过思考现象发现错误信号可线性解码
- 5类固定线性修正均无效果(Δ≈0)
- 错误结构可用于事后可靠性判断
大模型隐藏状态中的错误信号是否可被线性解码并用于纠正?我们通过过思考(OT)这一稳定行为模式(杰卡德指数≥0.81,标注者一致性94%)进行研究:模型在重采样下正确回答,但在扩展思维链中失败。该现象在线性解码下达到71.6%平衡准确率(p < 10^{-16})。然而,五类固定线性转向(29配置,n=1,273)均未产生效果(Δ≈0),跨架构(Qwen2.5-7B)与跨领域(MMLU-STEM)结果一致。三组证据表明表征纠缠:OT方向与任务关键计算有85-88%重叠(特异性比≤0.152);非目标共享方向修正导致准确率下降12.1个百分点;LEACE概念擦除使准确率下降3.6个百分点(p=0.01),而10次随机擦除仅提升0.3个百分点。实例级探针-转向相关性为r=-0.002(p=0.97)。积极方面,同一探针可实现选择性拒答(保留集AUROC=0.610,优于所有五个不确定性基线,p=0.009):即使无法修正,解码的错误结构仍支持生成后可靠性估计。
原文摘要 · Abstract (English)
Can linearly decodable failure signals in LLM hidden states be leveraged to correct those failures? We investigate this classification-correction gap via Overthinking (OT)--a stable behavioral regime (Jaccard >= 0.81, 94% inter-annotator agreement) in medical QA where models answer correctly under resampling yet fail in extended chain-of-thought. OT is linearly decodable at 71.6% balanced accuracy (p < 10^{-16}). Yet five families of fixed linear steering (29 configurations, n=1,273) all yield Delta ~= 0, with identical null results cross-architecture (Qwen2.5-7B) and cross-domain (MMLU-STEM). Three convergent lines of evidence suggest representational entanglement: the OT direction has 85-88% overlap with task-critical computation (specificity ratio <= 0.152); non-targeted shared-direction steering damages accuracy (-12.1pp); and LEACE concept erasure damages accuracy (-3.6pp, p=0.01), while 10 random erasures produce Delta=+0.3pp. The per-instance probe-steering correlation is r=-0.002 (p=0.97). Positively, the same probe enables selective abstention (held-out AUROC=0.610, exceeding all five uncertainty baselines, p=0.009): decodable failure structure supports post-generation reliability estimation even when the fixed linear steering family cannot exploit it for correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。