arXiv:2607.29484cs.CLcs.LG2026-07

干预数据未必教会模型因果方向,上下文证据类型才是关键。

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

论文配图:Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
图 1 · 摘自论文原文
  • 用合成环境测试发现干预数据无法纠正错误因果判断
  • 93亿参数模型中仍有6%的因果方向判断错误
  • 适合研究大模型因果推理机制与上下文依赖问题

我们在一个完全可控的合成环境中检验了干预数据是否能有效教导模型因果推理。结果发现,在辛普森悖论场景下,增加预训练中干预样本比例并不能提升对因果方向的判断能力:模型对do-操作的响应幅度持续增长,但符号仍由观察性上下文决定。真正影响因果判断的是推理时上下文中的证据类型。在纯观察性上下文中,58%(29/50)的世界出现系统性符号反转;混合上下文为38%(19/50),而仅使用一致的干预探针时正确率达82%(41/50)。移除上下文中的观察性证据后,因果推断能力立即恢复(true_ratio = +0.56)。该抑制效应在不同训练种子下稳定存在(11/11次强反转),且在93亿参数模型中依然显著(匹配探针组中31.8%对6%的反转率)。外部审计揭示模型存在正向效应先验,两层结构可被随机重训消除分布内偏差,但无法消除分布外偏差。结论是:因果能力存在于权重中,而开关在上下文中,激活修补定位到中间层的观察行。我们还量化了探针评估的采样噪声,并提出一种证据平均协议,将符号错误从26%降至9%。

原文摘要 · Abstract (English)

Interventional data is widely regarded as the gold standard for teaching models causal reasoning. We test this assumption in a fully controlled synthetic environment pitting observational correlation against causal effect, and find it fails instructively. In Simpson's-paradox worlds, where the two have systematically opposite signs, increasing the fraction of interventional samples in pretraining does not improve causal direction: the magnitude of the model's do()-response grows monotonically, yet its sign is copied from the observational context. What governs whether interventional evidence is used is not the training mixture but the evidence type present in the context at inference time. Under an identical training recipe, a purely observational context induces systematic sign reversal in 29/50 worlds, a mixed context in 19/50, while aligned interventional probes alone yield 41/50 correct. Erasing observational evidence from the context immediately releases the suppressed causal interpolation ability (ratio_true = +0.56); a four-state content manipulation shows the switch is content-mediated and graded. The suppression is stable across training seeds (11/11 strong reversals persist on a matched-protocol second seed) and robust as a rate at 0.93B parameters (31.8% vs. 6% reversals in the matched probe-only arm), even as absolute gains shrink four-fold. An external audit on CLadder exposes a learned positive-effect prior with a two-layer structure: sign-randomized retraining removes it in-distribution but not out-of-distribution. We summarize: the capability lives in the weights; the switch lives in the context, and activation patching localizes the switch to the middle layers' observational rows. We further quantify the sampling noise floor of probe-based causal evaluation and an evidence-averaging protocol that cuts sign errors from 26% to 9%.

因果推理大模型上下文依赖干预数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。