大模型懂因果但说不清,输出常被常识误导
Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

- 用线性探测从隐藏状态提取真实因果答案
- 输出准确率仅约0.5,远低于隐藏状态的0.97
- 揭示模型内部理解与口头回答的断裂
我们发现大语言模型对因果问题的内部表征与其口头回答之间存在不匹配。在反常识的CLadder任务中,固定线性探测器能从模型隐藏状态中恢复出证据支持的答案(准确率约0.97),而其口头的“是/否”回答却回归常识判断(准确率约0.5)。这一约0.5的差距被称为因果舌结(Causal Tongue-Tie):错误的“是/否”可分解为两种独立失败模式——内部无信号,或有信号但无法表达。这对仅依赖输出的因果基准测试具有双重启示:回答正确未必代表理解,回答错误也未必代表无法推理。基于单一准确率得出的大模型是否具备因果推理能力的普遍结论,值得重新审视。
原文摘要 · Abstract (English)
We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed linear probe recovers the evidence-supported answer from the model's hidden state (accuracy approximately 0.97), while the spoken Yes/No reverts to the commonsense one (accuracy approximately 0.5). We call this approximately +0.5 gap Causal Tongue-Tie: a wrong Yes/No decomposes into two separable failure modes: no internal signal versus a signal the verbal interface cannot say. The implication cuts both ways for output-only causal benchmarks: a benchmark "correct" need not mean the model has understood, and a benchmark "wrong" need not mean it cannot. Sweeping claims about whether LLMs can do causal reasoning, drawn from a single accuracy number, deserve a second look.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。