提出新方法解决因果推断中的条件矩约束问题,显著降低深度模型估计偏差。
Double Machine Learning for Conditional Moment Restrictions: IV Regression, Proximal Causal Learning and Beyond
- 基于双重机器学习框架设计新算法,用新目标函数减小估计偏差。
- 在真实数据上表现超越现有方法,达到最优收敛速度 $O(N^{-1/2})$。
- 适合使用深度神经网络做因果推断的研究者,尤其适用于工具变量和近似因果学习。
求解条件矩约束(CMR)是统计学、因果推断和计量经济学中的核心问题,目标是找到满足特定条件矩等式的函数。许多因果推断技术,如工具变量(IV)回归和近似因果学习(PCL),都属于CMR问题。现有方法多采用两阶段估计,将第一阶段结果直接代入第二阶段,但直接代入会导致第二阶段严重偏差,尤其在两阶段均使用深度神经网络(DNN)时,正则化与过拟合会加剧偏差。本文提出DML-CMR,一种基于双重/去偏机器学习框架的两阶段CMR估计器,通过设计新型学习目标减少偏差。理论证明其在参数化假设和弱正则性条件下可实现最小最大最优收敛率 $O(N^{-1/2})$,其中 $N$ 为样本量。在真实数据集上,将DML-CMR应用于包含深度神经网络的IV回归与近似因果学习,性能优于现有CMR方法及针对特定问题定制的算法,达到当前最优水平。
原文摘要 · Abstract (English)
Solving conditional moment restrictions (CMRs) is a key problem considered in statistics, causal inference, and econometrics, where the aim is to solve for a function of interest that satisfies some conditional moment equalities. Specifically, many techniques for causal inference, such as instrumental variable (IV) regression and proximal causal learning (PCL), are CMR problems. Most CMR estimators use a two-stage approach, where the first-stage estimation is directly plugged into the second stage to estimate the function of interest. However, naively plugging in the first-stage estimator can cause heavy bias in the second stage. This is particularly the case for recently proposed CMR estimators that use deep neural network (DNN) estimators for both stages, where regularisation and overfitting bias is present. We propose DML-CMR, a two-stage CMR estimator that provides an unbiased estimate with fast convergence rate guarantees. We derive a novel learning objective to reduce bias and develop the DML-CMR algorithm following the double/debiased machine learning (DML) framework. We show that our DML-CMR estimator can achieve the minimax optimal convergence rate of $O(N^{-1/2})$ under parameterisation and mild regularity conditions, where $N$ is the sample size. We apply DML-CMR to a range of problems using DNN estimators, including IV regression and proximal causal learning on real-world datasets, demonstrating state-of-the-art performance against existing CMR estimators and algorithms tailored to those problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。