arXiv:2602.06899cs.LGstat.ML2026-02

融合时间与环境异质性,提升因果图可识别性与统计可恢复性。

Sample Complexity of Causal Identification with Temporal Heterogeneity

  • 结合时间动态与多环境异质性,构建统一可识别条件。
  • 在重尾噪声下样本复杂度显著上升,代价由信息论边界量化。
  • 适用于非平稳系统中因果推断的实用场景,尤其关注数据稀缺时

从观测数据中恢复唯一因果图是病态问题,因多种生成机制可产生相同观测分布。仅通过引入特定结构或分布假设才能使问题可解。现有研究分别利用时间序列动态或多环境异质性来约束该问题,本文将二者作为互补异质性来源进行整合,推导出统一的必要可识别条件,并对轻尾与重尾噪声下的统计恢复极限进行了严格分析。结果表明,时间结构可有效替代缺失的环境多样性,即使在异质性不足时仍可能实现可识别性。进一步扩展至重尾(Student's t)分布时,虽几何可识别条件不变,但样本复杂度相比高斯基准显著发散。显式的信息论界量化了鲁棒性的代价,确立了协方差基础因果图恢复方法在真实非平稳系统中的基本极限。本工作将关注点从‘是否可识别’转向‘是否可统计恢复’。

原文摘要 · Abstract (English)

Recovering a unique causal graph from observational data is an ill-posed problem because multiple generating mechanisms can lead to the same observational distribution. This problem becomes solvable only by exploiting specific structural or distributional assumptions. While recent work has separately utilized time-series dynamics or multi-environment heterogeneity to constrain this problem, we integrate both as complementary sources of heterogeneity. This integration yields unified necessary identifiability conditions and enables a rigorous analysis of the statistical limits of recovery under thin versus heavy-tailed noise. In particular, temporal structure is shown to effectively substitute for missing environmental diversity, possibly achieving identifiability even under insufficient heterogeneity. Extending this analysis to heavy-tailed (Student's t) distributions, we demonstrate that while geometric identifiability conditions remain invariant, the sample complexity diverges significantly from the Gaussian baseline. Explicit information-theoretic bounds quantify this cost of robustness, establishing the fundamental limits of covariance-based causal graph recovery methods in realistic non-stationary systems. This work shifts the focus from whether causal structure is identifiable to whether it is statistically recoverable in practice.

因果推断样本复杂度时间序列重尾分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。