arXiv:2502.12323cs.LGstat.ML2025-02

用对抗学习修正模型预测偏差,让回归分析更准确。

Adversarial Debiasing for Unbiased Parameter Recovery

  • 引入对抗训练使预测结果去偏
  • 在非洲森林覆盖数据上恢复真实系数
  • 适合用机器学习预测做因变量的社科研究者

机器学习模型的预测误差会引发回归系数估计偏差,影响社会科学中以模型预测为因变量的研究。本文揭示了偏差来源,提出检测方法,并运用对抗学习算法对预测结果去偏。该方法适用于任何将机器学习预测作为回归因变量的场景。通过模拟及基于非洲森林覆盖的真实卫星数据实验发现,未经校正的模型预测会导致参数估计偏差,而使用对抗训练后的模型能准确恢复真实系数。

原文摘要 · Abstract (English)

Advances in machine learning and the increasing availability of high-dimensional data have led to the proliferation of social science research that uses the predictions of machine learning models as proxies for measures of human activity or environmental outcomes. However, prediction errors from machine learning models can lead to bias in the estimates of regression coefficients. In this paper, we show how this bias can arise, propose a test for detecting bias, and demonstrate the use of an adversarial machine learning algorithm in order to de-bias predictions. These methods are applicable to any setting where machine-learned predictions are the dependent variable in a regression. We conduct simulations and empirical exercises using ground truth and satellite data on forest cover in Africa. Using the predictions from a naive machine learning model leads to biased parameter estimates, while the predictions from the adversarial model recover the true coefficients.

去偏对抗学习回归分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。