arXiv:2510.25128cs.LGstat.ML2025-10NeurIPS被引 3

数据增强可缓解隐藏混杂带来的因果效应估计偏差。

An Analysis of Causal Effect Estimation using Outcome Invariant Data Augmentation

  • 将数据增强视为对处理机制的干预,利用结果不变性提升因果推断性能。
  • 提出类工具变量回归方法,在弱化传统工具变量条件时仍有效减少偏倚。
  • 实验证明该方法在真实数据与模拟数据中均优于普通数据增强策略。

数据增强通常用于独立同分布设置下的正则化以提升泛化能力。本文提出一个统一框架,将数据增强与因果推断结合,论证其不仅适用于 i.i.d. 场景,还能提升对干预的泛化能力。当结果生成机制对数据增强保持不变时,此类增强可被视为对处理生成机制的干预,从而有助于缓解由隐藏混杂引起的因果效应估计偏差。在存在未观测混杂时,通常依赖工具变量(IV),但其获取困难。为此,本文通过正则化基于 IV 的估计器,引入类工具变量(IVL)回归,在放松部分 IV 假设条件下仍能减轻混杂偏倚并提升跨干预预测性能。进一步将参数化数据增强建模为 IVL 回归问题,证明其组合使用可模拟最坏情况下的数据增强,显著提升因果估计与泛化任务表现。理论分析覆盖总体情形,仿真实验基于线性模型验证有限样本表现,真实数据实验也支持核心观点。

原文摘要 · Abstract (English)

The technique of data augmentation (DA) is often used in machine learning for regularization purposes to better generalize under i.i.d. settings. In this work, we present a unifying framework with topics in causal inference to make a case for the use of DA beyond just the i.i.d. setting, but for generalization across interventions as well. Specifically, we argue that when the outcome generating mechanism is invariant to our choice of DA, then such augmentations can effectively be thought of as interventions on the treatment generating mechanism itself. This can potentially help to reduce bias in causal effect estimation arising from hidden confounders. In the presence of such unobserved confounding we typically make use of instrumental variables (IVs) -- sources of treatment randomization that are conditionally independent of the outcome. However, IVs may not be as readily available as DA for many applications, which is the main motivation behind this work. By appropriately regularizing IV based estimators, we introduce the concept of IV-like (IVL) regression for mitigating confounding bias and improving predictive performance across interventions even when certain IV properties are relaxed. Finally, we cast parameterized DA as an IVL regression problem and show that when used in composition can simulate a worst-case application of such DA, further improving performance on causal estimation and generalization tasks beyond what simple DA may offer. This is shown both theoretically for the population case and via simulation experiments for the finite sample case using a simple linear example. We also present real data experiments to support our case.

因果推断数据增强工具变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。