arXiv:2505.13422econ.EMcs.LG2025-05被引 1

用机器学习做2SLS第一阶段,线性方法有效,非线性可能更糟。

Machine learning the first stage in 2SLS: Practical guidance from bias decomposition and simulation

  • 将线性机器学习(如后Lasso)用于2SLS第一阶段,可降低偏差。
  • 非线性方法(如随机森林、神经网络)会使第二阶段估计偏差剧增。
  • 适合做因果推断的学者参考,避免误用非线性模型。

机器学习主要针对预测问题发展而来,而两阶段最小二乘法(2SLS)的第一阶段本质上是预测问题,暗示使用机器学习可能带来收益。然而,目前缺乏关于何时机器学习能提升2SLS、何时反而有害的实用指导。本文通过偏差分解,揭示了将机器学习引入2SLS的三类关键偏差来源。机制上,这类方法同时面临预测与因果推断场景的共性问题及其交互影响。通过模拟实验,发现线性机器学习方法(如后Lasso)表现良好,而非线性方法(如随机森林、神经网络)会导致第二阶段估计产生显著偏差,甚至超过内生性普通最小二乘法(OLS)的偏差。

原文摘要 · Abstract (English)

Machine learning (ML) primarily evolved to solve "prediction problems." The first stage of two-stage least squares (2SLS) is a prediction problem, suggesting potential gains from ML first-stage assistance. However, little guidance exists on when ML helps 2SLS$\unicode{x2014}$or when it hurts. We investigate the implications of inserting ML into 2SLS, decomposing the bias into three informative components. Mechanically, ML-in-2SLS procedures face issues common to prediction and causal-inference settings$\unicode{x2014}$and their interaction. Through simulation, we show linear ML methods (e.g., post-Lasso) work well, while nonlinear methods (e.g., random forests, neural nets) generate substantial bias in second-stage estimates$\unicode{x2014}$potentially exceeding the bias of endogenous OLS.

因果推断机器学习2SLS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。