arXiv:2411.01295cs.LGcs.AI2024-11NeurIPS被引 10

用流模型生成可验证因果推断的合成数据,精确控制处理效应。

Marginal Causal Flows for Validation and Inference

  • 基于归一化流构建灵活数据生成模型,直接从观测数据推断因果量。
  • 生成的合成数据与真实数据相似,且平均处理效应精确符合设定值。
  • 适合验证因果推断方法,尤其适用于需要精确控制混杂的场景。

由于现有模型灵活性不足以及因果基准数据集复杂度有限,从复杂数据中研究干预对结果的边际因果效应仍具挑战。本文提出 Frugal Flows,一种基于似然的新型机器学习模型,利用归一化流灵活学习数据生成过程,并直接从观测数据中推断边际因果量。该模型特别适合生成用于验证因果方法的合成数据:既能紧密模仿真实数据分布,又能自动且精确满足用户定义的平均处理效应(ATE)。据我们所知,Frugal Flows 是首个同时具备灵活数据建模能力与精确参数化因果量(如 ATE 及未观测混杂程度)的生成模型。我们在模拟和真实数据集上进行了实验验证。

原文摘要 · Abstract (English)

Investigating the marginal causal effect of an intervention on an outcome from complex data remains challenging due to the inflexibility of employed models and the lack of complexity in causal benchmark datasets, which often fail to reproduce intricate real-world data patterns. In this paper we introduce Frugal Flows, a novel likelihood-based machine learning model that uses normalising flows to flexibly learn the data-generating process, while also directly inferring the marginal causal quantities from observational data. We propose that these models are exceptionally well suited for generating synthetic data to validate causal methods. They can create synthetic datasets that closely resemble the empirical dataset, while automatically and exactly satisfying a user-defined average treatment effect. To our knowledge, Frugal Flows are the first generative model to both learn flexible data representations and also exactly parameterise quantities such as the average treatment effect and the degree of unobserved confounding. We demonstrate the above with experiments on both simulated and real-world datasets.

因果推断生成模型合成数据归一化流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。