通过非凸优化实现因果不变性学习,高效识别变量间真实因果关系。
Causal Invariance Learning via Efficient Nonconvex Optimization
- 基于多环境数据的不变性原理,设计负权分布鲁棒优化框架。
- 理论证明所有稳定点都接近真实因果模型,且算法可收敛到近似最优解。
- 避免指数级特征子集搜索,适用于高维数据,适合因果推断研究者。
从观测数据中识别变量间的因果关系是一项重要但具有挑战性的任务。本文聚焦于识别结果的直接原因并估计其影响大小,即学习因果结果模型。来自多个异质环境的数据可通过不变性原理提供揭示因果关系的机会,即因果结果模型在不同环境中保持不变。基于此,我们提出负权重分布鲁棒优化(NegDRO)框架,以学习不变预测模型。NegDRO通过最小化跨环境最坏情况风险组合并允许负权重来强制实现不变性。在加性干预设定下,本文有三大贡献:(i) 统计方面,给出充分且几乎必要的识别条件,使不变预测模型与因果结果模型一致;(ii) 优化方面,尽管NegDRO是非凸的,但其优化景观表现良好,所有驻点均接近真实因果模型;(iii) 计算方面,提出基于梯度的算法,能保证收敛至因果模型,且给出了样本量与梯度迭代次数的非渐近收敛速率。特别地,该方法避免了文献中对指数级协变量子集的穷举搜索,确保高维情形下的可扩展性。据我们所知,这是首个能够高效求解非凸优化问题并找到近似全局最优解的因果不变性学习方法。
原文摘要 · Abstract (English)
Identifying the causal relationship among variables from observational data is an important yet challenging task. This work focuses on identifying the direct causes of an outcome and estimating their magnitude, i.e., learning the causal outcome model. Data from multiple environments provide valuable opportunities to uncover causality by exploiting the invariance principle that the causal outcome model holds across heterogeneous environments. Based on the invariance principle, we propose the Negative Weighted Distributionally Robust Optimization (NegDRO) framework to learn an invariant prediction model. NegDRO minimizes the worst-case combination of risks across multiple environments and enforces invariance by allowing potential negative weights. Under the additive interventions regime, we establish three major contributions: (i) On the statistical side, we provide sufficient and nearly necessary identification conditions under which the invariant prediction model coincides with the causal outcome model; (ii) On the optimization side, despite the nonconvexity of NegDRO, we establish its benign optimization landscape, where all stationary points lie close to the true causal outcome model; (iii) On the computational side, we develop a gradient-based algorithm that provably converges to the causal outcome model, with non-asymptotic convergence rates in both sample size and gradient-descent iterations. In particular, our method avoids exhaustive combinatorial searches over exponentially many subsets of covariates found in the literature, ensuring scalability even when the dimension of the covariates is large. To our knowledge, this is the first causal invariance learning method that finds the approximate global optimality for a nonconvex optimization problem efficiently.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。