用预训练文本嵌入提升因果推断,处理文本型混杂变量
From Text to Treatment Effects: A Meta-Learning Approach to Handling Text-Based Confounding
- 用预训练文本嵌入+表格变量联合建模混杂因素
- 数据充足时,文本信息显著提升处理效应估计精度
- 适合研究文本数据中因果关系的学者参考
因果机器学习的核心目标之一是从观测数据中准确估计异质处理效应。近年来,元学习成为一种灵活、模型无关的框架,可利用任意监督模型估计条件平均处理效应(CATE)。本文探讨了当混杂变量以文本形式存在时,元学习方法的表现。通过合成数据实验发现,在使用预训练文本表示作为混杂变量补充的同时,结合表格背景变量,相比仅依赖表格变量的模型,能获得更优的CATE估计,尤其在数据量充足时表现更佳。然而,由于文本嵌入存在语义纠缠,模型性能仍无法达到具备完整混杂变量知识的元学习器水平。该结果揭示了预训练文本表示在因果推断中的潜力与局限,也为未来研究提供了新方向。
原文摘要 · Abstract (English)
One of the central goals of causal machine learning is the accurate estimation of heterogeneous treatment effects from observational data. In recent years, meta-learning has emerged as a flexible, model-agnostic paradigm for estimating conditional average treatment effects (CATE) using any supervised model. This paper examines the performance of meta-learners when the confounding variables are expressed in text. Through synthetic data experiments, we show that learners using pre-trained text representations of confounders, in addition to tabular background variables, achieve improved CATE estimates compared to those relying solely on the tabular variables, particularly when sufficient data is available. However, due to the entangled nature of the text embeddings, these models do not fully match the performance of meta-learners with perfect confounder knowledge. These findings highlight both the potential and the limitations of pre-trained text representations for causal inference and open up interesting avenues for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。