arXiv:2603.17384cs.LG2026-03被引 1

用拓扑学揭示生成模型的反事实缺陷,提出可计算的因果建模新框架。

Cohomological Obstructions to Global Counterfactuals: A Sheaf-Theoretic Foundation for Generative Causal Models

  • 将因果模型建模为水街空间上的层结构,用上同调障碍分析全局反事实矛盾。
  • 引入熵正则化,构建可微分的因果层拉普拉斯算子,实现零内存反向传播。
  • 适用于高维单细胞数据等存在拓扑障碍的因果发现场景。

当前连续生成模型(如扩散模型、流匹配)隐含假设:局部一致的因果机制自然产生全局一致的反事实。本文证明,当因果图具有非平凡同调结构(如结构冲突或隐藏混杂因子)时,该假设根本性失效。我们将结构因果模型形式化为定义在 Wasserstein 空间上的细胞层,严格给出测度空间中上同调障碍的代数拓扑定义。为保证计算可行性并避免确定性奇点(我们定义为流形撕裂),引入熵正则化,推导出熵正则化的 Wasserstein 因果层拉普拉斯算子——一种耦合的非线性 Fokker-Planck 方程系统。关键地,我们证明了推送测度一阶变分的熵拉回引理。结合 Sinkhorn 最优性条件上的隐函数定理,建立了与自动微分(VJP)的直接算法桥梁,实现严格与迭代步数无关的 O(1) 内存反向梯度。实验表明,该框架成功利用热噪声穿越高维 scRNA-seq 反事实中的拓扑屏障(“熵隧穿”)。最后,逆向使用该理论提出拓扑因果得分,证明我们的层拉普拉斯算子可作为高度敏感的拓扑感知因果发现探测器。

原文摘要 · Abstract (English)

Current continuous generative models (e.g., Diffusion Models, Flow Matching) implicitly assume that locally consistent causal mechanisms naturally yield globally coherent counterfactuals. In this paper, we prove that this assumption fails fundamentally when the causal graph exhibits non-trivial homology (e.g., structural conflicts or hidden confounders). We formalize structural causal models as cellular sheaves over Wasserstein spaces, providing a strict algebraic topological definition of cohomological obstructions in measure spaces. To ensure computational tractability and avoid deterministic singularities (which we define as manifold tearing), we introduce entropic regularization and derive the Entropic Wasserstein Causal Sheaf Laplacian, a novel system of coupled non-linear Fokker-Planck equations. Crucially, we prove an entropic pullback lemma for the first variation of pushforward measures. By integrating this with the Implicit Function Theorem (IFT) on Sinkhorn optimality conditions, we establish a direct algorithmic bridge to automatic differentiation (VJP), achieving O(1)-memory reverse-mode gradients strictly independent of the iteration horizon. Empirically, our framework successfully leverages thermodynamic noise to navigate topological barriers ("entropic tunneling") in high-dimensional scRNA-seq counterfactuals. Finally, we invert this theoretical framework to introduce the Topological Causal Score, demonstrating that our Sheaf Laplacian acts as a highly sensitive algebraic detector for topology-aware causal discovery.

因果模型拓扑学习生成模型熵正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。