arXiv:2606.15733cs.CLcs.AI2026-06

发现语言模型因果推理因变量名变化而错答,根源是表征对齐问题。

Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning

  • 用双视图权重更新探测表征对齐机制
  • 替换变量名后模型决策仍可保留答案信息
  • 适合研究大模型推理偏差与表征对齐的学者

指令微调的语言模型在英文变量名被同类型占位符替换后,会以不同方式回答相同的因果推理问题,尽管结构因果模型和正确答案未变。我们探究这一词汇间隙是源于占位符视图中的信息丢失,还是仍携带答案相关内容的表征存在读出错位。Vernier 采用配对视图权重更新作为工具,分析间隙消除后的机制。在有效设置下,证据支持表征错位。变量名探针在占位符视图上的准确率提升,且对 Qwen-7B、Qwen-14B、Llama-3.1-8B 的激活修补显示,决策词的表征可在两视图间传递答案身份。使视图对齐的更新是原始与占位符提示的反事实增强,而答案子空间 KL 主要强化中间答案信念的一致性。成功受限于模型家族、规模和任务:CRASS 跨 Qwen 规模可靠,e-CARE 表现弱,初步非因果重命名任务亦呈现类似定性模式。

原文摘要 · Abstract (English)

Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced by type-preserving placeholders, although the structural causal model and the gold answer are unchanged. We ask whether this lexical gap reflects information loss in the placeholder view or a misaligned read-out from a representation that still carries answer-relevant content. Vernier uses a paired-view weight update as an instrument and then inspects the mechanism left after the gap closes. In the working regimes, the evidence favours representational misalignment. A variable-name probe becomes more accurate on the placeholder view, and activation patching on Qwen-7B, Qwen-14B, and Llama-3.1-8B shows that the decision-token representation can transfer answer identity between views. The update that realigns the views is counterfactual augmentation over original and placeholder prompts, while the answer-subspace KL mainly sharpens intermediate answer-belief agreement. Success is bounded by model family, scale, and task. CRASS transfer is reliable across Qwen scales and Llama, e-CARE remains weak, and preliminary non-causal rename tasks show a similar qualitative pattern.

因果推理表征对齐大模型偏差语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。