arXiv:2605.15353cs.LGcs.AI2026-05中稿 · ICML

基于干预数据的因果图发现,通过构造保证无环,大幅提升效率与稳定性。

PACER: Acyclic Causal Discovery from Large-Scale Interventional Data

论文配图:PACER: Acyclic Causal Discovery from Large-Scale Interventional Data
图 1 · 摘自论文原文
  • 用变量排列与边概率联合建模,直接搜索有效因果结构
  • 在蛋白质信号和基因扰动数据上性能达或超越现有方法
  • 支持大规模数据,速度比传统方法快100倍以上

从数据中推断有向无环图(DAG)是因果发现的核心挑战,尤其在高维、大规模干预数据日益丰富的背景下。现有方法受限于软无环约束,导致优化过程涉及无效循环图,引发数值不稳和可扩展性差的问题。本文提出PACER(扰动驱动的无环因果边恢复),一种可扩展的因果发现框架,通过构造确保无环性。该框架通过变量排列与边概率的联合模型对DAG分布进行参数化,实现对合法因果结构的直接优化,无需替代惩罚项。其统一处理观测与干预数据,支持灵活的条件密度模型及结构先验知识。针对线性高斯机制,推导出干预对数似然及其梯度的闭式表达,带来显著计算优势。实验表明,PACER在蛋白质信号和大规模基因扰动基准上表现达到或超过当前最优,可高效处理数千变量网络,相较基于惩罚项的可微方法提速高达两个数量级。结果证明,通过合理设计搜索空间,高维扰动数据下的精确且可扩展的因果发现是可行的。

原文摘要 · Abstract (English)

Inferring the structure of directed acyclic graphs (DAGs) from data is a central challenge in causal discovery, particularly in modern high-dimensional settings where large-scale interventional data are increasingly available. While interventional data can improve identifiability, existing methods remain limited by soft acyclicity constraints, leading to optimization over invalid cyclic graphs, numerical instability, and reduced scalability. We introduce PACER (Perturbation-driven Acyclic Causal Edge Recovery), a scalable framework for causal discovery that guarantees acyclicity by construction. PACER parameterizes a distribution over DAGs through a joint model of variable permutations and edge probabilities, enabling direct optimization over valid causal structures without surrogate penalties. The framework supports a unified likelihood-based treatment of observational and interventional data, flexible conditional density models, and the incorporation of structural prior knowledge. For linear-Gaussian mechanisms, we derive closed-form expressions for the expected interventional log-likelihood and its gradients, yielding substantial computational gains. Empirically, PACER matches or exceeds state-of-the-art methods on protein signaling and large-scale genetic perturbation benchmarks, while scaling efficiently to networks with thousands of variables and achieving up to two orders of magnitude speedups over penalty-based differentiable approaches. These results demonstrate that exact and scalable causal discovery from high-dimensional perturbation data is achievable through principled search space design.

因果发现无环图干预数据可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。