arXiv:2606.18074stat.MLcs.LG2026-06

用协方差张量实现高效因果发现,可处理数百变量

Tensor-based second-order causal discovery

  • 基于观测与干预数据的协方差张量构造因果模型
  • 仅需对数级干预次数即可识别因果顺序与参数
  • 适用于线性与非线性模型,对噪声鲁棒且可扩展

因果发现旨在揭示变量间的因果依赖关系。为此,我们提出一种名为张量化二阶因果发现(TSCD)的算法,其输入为观测与干预数据协方差矩阵构成的张量。在假设因果依赖服从有向无环图(DAG)上的线性结构方程模型的前提下,TSCD 能输出 DAG 及其边上的函数,仅需噪声变量互不相关。我们还实现了该方法在非线性模型中的版本。聚焦二阶统计量(通过协方差矩阵)是因其相较于高阶矩具有更高的统计与计算效率,且相比一阶统计量具备可辨识性,且不依赖变量是否服从高斯分布。我们证明,TSCD 可在干预次数仅为变量数的对数级别时实现因果顺序与参数的可辨识。实验表明,TSCD 对噪声具有鲁棒性,在性能上与现有方法相当,并可扩展至数百变量。

原文摘要 · Abstract (English)

Causal discovery seeks to uncover the causal dependencies among variables. For this purpose, we propose an algorithm called Tensor-based Second-order Causal Discovery (TSCD). Its input is a tensor obtained from the covariance matrices of observational and interventional data. Assuming the causal dependencies follow a linear structural equation model on a directed acyclic graph (DAG), TSCD outputs the DAG and the functions on its edges, requiring only that the noise variables are uncorrelated. We also implement a version of the approach for nonlinear models. Our focus on second-order statistics (via the covariance matrices) is motivated by their statistical and computational efficiency relative to higher-order moments, their identifiability relative to first-order statistics, and that they work regardless of whether the variables are Gaussian. We show that TSCD has identifiable causal order and parameters from a number of interventions that is logarithmic in the number of variables. Experiments show that TSCD is robust to noise, competitive with existing methods, and scales to hundreds of variables.

因果发现张量方法二阶统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。