提出MAR-S框架,让神经网络预测结果更可靠、可复现且无需过度优化。
A Unifying Framework for Robust and Efficient Inference with Unstructured Data
- 用验证样本修正神经网络预测误差,实现无偏推断
- 支持描述性与因果性估计,适配聚合和变换后的预测结果
- 解决模型选择自由度高、结果难复现等问题,适合实证研究者
为分析非结构化数据(文本、图像、音频、视频),经济学家通常先用神经网络提取低维结构特征。但神经网络本身存在系统偏差,这些偏差会传递至后续估计量。传统做法将提取的结构变量视为代理变量,隐含接受任意测量误差,而当前快速演进的AI使数据提取成本极低,由此引发一系列问题:研究者自由度高(如模型架构、训练数据或提示词选择等)导致p值操纵风险,专有模型难以复现,且缺乏判断预测精度是否足够高的标准。为此,本文提出MAR-S(Missing At Random Structured Data)——一种半参数缺失数据框架,通过验证样本校正神经网络预测误差,实现无偏、高效且鲁棒的推断。MAR-S整合并扩展现有基于机器学习预测的去偏方法,将其与因果推断等经典问题联系起来,开发出适用于描述性与因果性目标的稳健高效估计器,并处理聚合与变换后预测结果的推断问题,填补了现有文献空白。
原文摘要 · Abstract (English)
To analyze unstructured data (text, images, audio, video), economists typically first extract low-dimensional structured features with a neural network. Neural networks do not make generically unbiased predictions, and biases will propagate to estimators that use their predictions. While structured variables extracted from unstructured data have traditionally been treated as proxies - implicitly accepting arbitrary measurement error - this poses various challenges in an era where constantly evolving AI can cheaply extract data. Researcher degrees of freedom (e.g., the choice of neural network architecture, training data or prompts, and numerous implementation details) raise concerns about p-hacking and how to best show robustness, the frequent deprecation of proprietary neural networks complicates reproducibility, and researchers need a principled way to determine how accurate predictions need to be before making costly investments to improve them. To address these challenges, this study develops MAR-S (Missing At Random Structured Data), a semiparametric missing data framework that enables unbiased, efficient, and robust inference with unstructured data, by correcting for neural network prediction error with a validation sample. MAR-S synthesizes and extends existing methods for debiased inference using machine learning predictions and connects them to familiar problems such as causal inference, highlighting valuable parallels. We develop robust and efficient estimators for both descriptive and causal estimands and address inference with aggregated and transformed neural network predictions, a common scenario outside the existing literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。