用数据对称性统一多种因果表示学习方法,提升实际应用效果。
Unifying Causal Representation Learning with the Invariance Principle
- 以数据对称性替代严格因果假设,指导潜在变量识别
- 新方法在真实高维生态数据上显著改善治疗效应估计
- 适合需要灵活建模、不依赖强因果假设的研究者
因果表示学习(CRL)旨在从高维观测中恢复潜在因果变量,以解决干预预测或鲁棒分类等下游任务。现有方法针对不同问题设定,产生多种可辨识性结果,常被认为对应珍珠因果层级的不同层次,但这种对应并不总是精确。本文指出,许多方法实质上是通过匹配数据内在对称性来构建表示,而这些对称性未必具有因果意义。因此,我们提出一种统一框架,可混合使用不同假设(包括非因果假设),基于与问题相关的不变性原则进行选择。该方法显著提升了在真实高维生态数据上的治疗效应估计性能,澄清了因果假设在变量发现中的作用,并将研究重点转向保持数据对称性。
原文摘要 · Abstract (English)
Causal representation learning (CRL) aims at recovering latent causal variables from high-dimensional observations to solve causal downstream tasks, such as predicting the effect of new interventions or more robust classification. A plethora of methods have been developed, each tackling carefully crafted problem settings that lead to different types of identifiability. These different settings are widely assumed to be important because they are often linked to different rungs of Pearl's causal hierarchy, even though this correspondence is not always exact. This work shows that instead of strictly conforming to this hierarchical mapping, many causal representation learning approaches methodologically align their representations with inherent data symmetries. Identification of causal variables is guided by invariance principles that are not necessarily causal. This result allows us to unify many existing approaches in a single method that can mix and match different assumptions, including non-causal ones, based on the invariance relevant to the problem at hand. It also significantly benefits applicability, which we demonstrate by improving treatment effect estimation on real-world high-dimensional ecological data. Overall, this paper clarifies the role of causal assumptions in the discovery of causal variables and shifts the focus to preserving data symmetries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。