arXiv:2605.05029cs.LG2026-05

预测模型常学错因果关系,反而关注环境干扰。

The Predictive-Causal Gap: An Impossibility Theorem and Large-Scale Neural Evidence

论文配图:The Predictive-Causal Gap: An Impossibility Theorem and Large-Scale Neural Evidence
图 1 · 摘自论文原文
  • 在2695种配置中,最优编码器更关注环境而非系统本身
  • 高维场景下因果保真度降至10^-8,预测误差却低92%
  • 不加约束的模型在55%任务中误学环境主导表征

我们报告了预测表示学习中的系统性失败模式。在2695个训练用于预测线性高斯动力系统的神经网络配置中,最优编码器追踪的是环境而非目标系统。平均因果保真度(编码器对系统自由度的敏感度占比)为0.49,仅2.5%的配置超过0.70。该缺陷随维度增加而加剧:当N=100时,最优编码器变为因果盲(保真度~10^-8),但预测误差比因果表示低92%。我们证明此非优化偏差,而是预测目标的结构性缺陷——当环境模式较慢或噪声较小时,所有总体风险最小化解均编码环境。具有该预测-因果差距的动力系统在参数空间中构成开集且测度为正。在非线性杜芬-GRU实验中,无约束预测器在55%任务中学习环境主导表征(95%置信区间41–68%),而操作性接地条件下仅为24%(p=2.3e-3);在环境变化下,无约束模型的分布外均方误差膨胀达1.82倍,而接地模型仅为1.00倍。操作性接地可部分抑制该差距,但无显式系统-环境边界则无法恢复因果保真度。结果表明,预测-因果差距是学习的结构性限制,对自监督表示学习、世界模型及扩展范式有重要影响。

原文摘要 · Abstract (English)

We report a systematic failure mode in predictive representation learning. Across 2695 neural network configurations trained to predict linear-Gaussian dynamics, the optimal encoder tracks the environment rather than the system it is meant to model. The mean causal fidelity -- the fraction of encoder sensitivity allocated to system degrees of freedom -- is 0.49, and only 2.5% of configurations exceed 0.70. The failure intensifies with dimension: at N=100, the optimal encoder becomes causally blind (fidelity ~10^{-8}) while achieving 92% lower prediction error than the causal representation. We prove this is not an optimization artifact but a structural property of the predictive objective: when environment modes are slower or less noisy than system modes, every minimizer of the population risk encodes the former. The set of dynamics exhibiting this predictive-causal gap is open and of positive measure in parameter space. In a nonlinear Duffing-GRU sweep, unconstrained predictors learn environment-dominant representations in 55% of tasks (95% CI 41--68%) versus 24% under operational grounding (p=2.3e-3); the median out-of-distribution MSE inflation under environment shift is 1.82x versus 1.00x. Operational grounding -- restricting the loss to system observables -- partially suppresses the gap, but causal fidelity is never recovered without an explicit system-environment boundary. The results identify the predictive-causal gap as a structural limit of learning, with implications for self-supervised representation learning, world models, and the scaling paradigm.

因果学习表示学习世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。