用预训练模型特征调整混杂因素,提升非表格数据的因果推断准确率
Adjustment for Confounding using Pre-Trained Representations
- 利用预训练网络提取的潜在特征进行混杂因素调整
- 证明神经网络在高维特征下仍能实现快速收敛
- 适合做图像/文本等非表格数据的因果分析研究者
当前越来越多研究关注将平均处理效应(ATE)估计扩展至图像、文本等非表格数据,这些数据可能成为混杂因素来源。忽略其影响会导致结果偏差和错误科学结论。然而,处理非表格数据需复杂特征提取器,常结合迁移学习思想。本文研究如何利用预训练神经网络的潜在特征来调整混杂因素。我们形式化了这些特征实现有效调整与统计推断的条件,并以双重机器学习为例展示结果。讨论了潜在特征学习及下游参数估计中的关键挑战,包括高维性与表征不可识别性。传统针对加性或稀疏线性模型的结构假设在潜在特征中不现实。但研究表明,神经网络对这些问题具有鲁棒性,可通过适应学习问题的内在稀疏性与维度,实现快速收敛速率。
原文摘要 · Abstract (English)
There is growing interest in extending average treatment effect (ATE) estimation to incorporate non-tabular data, such as images and text, which may act as sources of confounding. Neglecting these effects risks biased results and flawed scientific conclusions. However, incorporating non-tabular data necessitates sophisticated feature extractors, often in combination with ideas of transfer learning. In this work, we investigate how latent features from pre-trained neural networks can be leveraged to adjust for sources of confounding. We formalize conditions under which these latent features enable valid adjustment and statistical inference in ATE estimation, demonstrating results along the example of double machine learning. We discuss critical challenges inherent to latent feature learning and downstream parameter estimation arising from the high dimensionality and non-identifiability of representations. Common structural assumptions for obtaining fast convergence rates with additive or sparse linear models are shown to be unrealistic for latent features. We argue, however, that neural networks are largely insensitive to these issues. In particular, we show that neural networks can achieve fast convergence rates by adapting to intrinsic notions of sparsity and dimension of the learning problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。