预测瓶颈不能发现因果结构,反而暴露了方法的陷阱。
Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
- 用线性瓶颈或Lasso可达到相同甚至更好效果
- 干预数据带来的优势主要受样本量影响,非因果信号
- 研究提供可复用的验证基准,适合方法对比与消融实验
一个仅用于下一步预测的Mamba状态空间模型看似能通过简单读出 $S = |W_{out} W_{in}|$ 恢复格兰杰因果结构,早期实验显示该现象在多种架构间泛化且干预数据带来显著优势($p < 10^{-5}$)。本文将验证流程封装为可复用的伪造检验基准:包含标准合成生成器(VAR/Lorenz/CauseMe)、三种干预语义($do(X=c)$、软噪声、随机强制)、三个真实数据集的边来源卡片以及匹配规模的对照组。经过五阶段测试发现:(i) 简单线性瓶颈表现不差;(ii) 调参后的Lasso在合成CauseMe和洛伦兹-96(唯一有明确真值的真实基准)上优于瓶颈,而经典PCMCI与格兰杰检验构成紧致领先群;(iii) 所谓干预优势约60%源于样本量混淆,在标准$do(X=c)$干预下消失,仅残留在非标准随机强制方案中;(iv) 即使如此,该效应也在经典双变量格兰杰检验中重现且更显著——说明其本质为方法无关。最终保留的是对瓶颈特性的狭义刻画结果,而基准本身才是持久贡献。
原文摘要 · Abstract (English)
A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |W_{out} W_{in}|$, with early experiments suggesting the phenomenon generalized across architectures and benefited from interventional data at $p < 10^{-5}$. We package the protocol used to test that claim -- standardized synthetic generators (VAR/Lorenz/CauseMe-style), three intervention semantics ($do(X=c)$, soft-noise, random-forcing), edge-provenance cards on three real datasets, and size-matched control arms -- as a reusable falsification benchmark, and walk the claim through it in five stages. The method-level claim does not survive: (i) a plain linear bottleneck does as well or better; (ii) tuned Lasso beats the bottleneck on synthetic CauseMe-style benchmarks, and on Lorenz-96 (the only real benchmark with unambiguous ground truth) classical PCMCI and Granger lead a tight cluster in which the bottleneck trails; (iii) the headline intervention advantage is roughly 60% a sample-size confound, and the residual disappears under standard $do(X=c)$ interventions, surviving only under a non-standard random-forcing scheme; (iv) even that residual reproduces, with a larger effect, in classical bivariate Granger -- the effect is method-agnostic. What survives is a narrow characterization result; the benchmark is the lasting artifact, and each stage above is one of its control arms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。