arXiv:2411.19245stat.MLcs.AI2024-11被引 7

提出对比学习方法,让高维治疗数据自动分离因果相关与无关特征。

Contrastive representations of high-dimensional, structured treatments

  • 用对比学习构建高维治疗的表示,分离因果相关与无关因素。
  • 理论证明该表示可实现无偏因果效应估计,实证在合成与真实数据上验证。
  • 适合处理文本、视频等复杂治疗场景的因果推断研究者。

因果效应估计对决策至关重要。传统方法通常处理二元或连续型治疗,但在许多现实场景中,治疗是高维结构化对象,如文本、视频或音频。这给传统因果推断带来挑战。尽管利用不同治疗间的共享结构有助于泛化到未见治疗,但本文指出,盲目使用该结构可能导致因果估计偏差。为此,我们提出一种新颖的对比学习方法,用于学习高维治疗的表示,并证明其能识别潜在因果因子并剔除非因果相关因子。理论证明该表示可实现无偏因果效应估计,且在合成和真实世界数据集上进行了实证验证与基准测试。

原文摘要 · Abstract (English)

Estimating causal effects is vital for decision making. In standard causal effect estimation, treatments are usually binary- or continuous-valued. However, in many important real-world settings, treatments can be structured, high-dimensional objects, such as text, video, or audio. This provides a challenge to traditional causal effect estimation. While leveraging the shared structure across different treatments can help generalize to unseen treatments at test time, we show in this paper that using such structure blindly can lead to biased causal effect estimation. We address this challenge by devising a novel contrastive approach to learn a representation of the high-dimensional treatments, and prove that it identifies underlying causal factors and discards non-causally relevant factors. We prove that this treatment representation leads to unbiased estimates of the causal effect, and empirically validate and benchmark our results on synthetic and real-world datasets.

因果推断对比学习高维治疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。