arXiv:2508.20259cs.LG2025-08中稿 · CIKM 2025被引 1

用隐变量模型提升因果推断在缺失协变量下的稳定性。

Latent Variable Modeling for Robust Causal Effect Estimation

  • 将隐变量引入双重机器学习第二阶段,分离表征学习与隐变量推断。
  • 在合成与真实数据集上验证了对隐藏混杂因素的鲁棒性。
  • 适合处理存在未观测混杂因素的因果分析场景。

隐变量模型为观测数据中未观测因素的建模与推断提供了有力框架。在因果推断中,它们有助于缓解因缺失或未测量协变量带来的挑战。本文提出一种新框架,将隐变量建模融入双重机器学习(DML)范式,以在存在隐藏因素时实现稳健的因果效应估计。考虑两种情形:一是隐变量仅影响结果,二是可能同时影响处理和结果。为保证可计算性,仅在DML的第二阶段引入隐变量,从而分离表示学习与隐变量推断过程。通过在合成与真实世界数据集上的广泛实验,验证了该方法的有效性与鲁棒性。

原文摘要 · Abstract (English)

Latent variable models provide a powerful framework for incorporating and inferring unobserved factors in observational data. In causal inference, they help account for hidden factors influencing treatment or outcome, thereby addressing challenges posed by missing or unmeasured covariates. This paper proposes a new framework that integrates latent variable modeling into the double machine learning (DML) paradigm to enable robust causal effect estimation in the presence of such hidden factors. We consider two scenarios: one where a latent variable affects only the outcome, and another where it may influence both treatment and outcome. To ensure tractability, we incorporate latent variables only in the second stage of DML, separating representation learning from latent inference. We demonstrate the robustness and effectiveness of our method through extensive experiments on both synthetic and real-world datasets.

因果推断隐变量双重机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。