用大模型直接预测条件期望,提升因果推断效率
Using LLMs to Directly Guess Conditional Expectations Can Improve Efficiency in Causal Estimation
- 用大模型生成的历史数据预测替代传统嵌入特征
- 在小规模拍卖数据上显著提升估计精度
- 适合需要高效处理高维混杂变量的因果研究者
我们提出一种简单有效的策略,利用大语言模型(LLM)驱动的AI工具改进因果估计。在双重机器学习框架中,治疗对结果的因果效应估计精度依赖于条件期望函数的估计性能。我们发现,基于历史数据训练的生成模型所作的预测,能比仅依赖模型提取嵌入的方法更有效地提升这些估计器的表现。我们认为,生成模型所具备的历史知识和推理能力有助于缓解因果推断中的维度诅咒问题。我们在一个小型在线珠宝拍卖数据集上进行了案例研究,结果表明,将大模型生成的预测值作为协变量加入模型,可显著提高估计效率。
原文摘要 · Abstract (English)
We propose a simple yet effective use of LLM-powered AI tools to improve causal estimation. In double machine learning, the accuracy of causal estimates of the effect of a treatment on an outcome in the presence of a high-dimensional confounder depends on the performance of estimators of conditional expectation functions. We show that predictions made by generative models trained on historical data can be used to improve the performance of these estimators relative to approaches that solely rely on adjusting for embeddings extracted from these models. We argue that the historical knowledge and reasoning capacities associated with these generative models can help overcome curse-of-dimensionality problems in causal inference problems. We consider a case study using a small dataset of online jewelry auctions, and demonstrate that inclusion of LLM-generated guesses as predictors can improve efficiency in estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。