用工具变量指导表示学习,解决高维治疗变量识别难题
Learning Treatment Representations for Downstream Instrumental Variable Regression
- 在表示学习中显式引入工具变量,避免隐式正则化导致的遗漏变量偏差
- 理论证明该方法可识别出最优结果预测方向,实验效果优于传统两阶段方法
- 适合处理高维、非结构化治疗数据(如病历路径)的因果推断场景
传统工具变量(IV)估计器存在根本性限制:可处理的内生治疗变量数量受限于可用工具变量数。当治疗变量以高维非结构化形式呈现(如医院中的患者治疗路径描述)时,研究者通常先使用无监督降维技术学习低维治疗表示,再进行IV回归分析。我们发现,此类方法因表示学习中的隐式正则化可能产生显著遗漏变量偏差。本文提出一种新方法,在表示学习过程中显式融合工具变量信息以构建治疗表示。该框架适用于工具变量有限但治疗变量高维的场景。理论与实证均表明,基于此工具变量引导表示的IV模型能确保识别出优化结果预测的方向。实验显示,该方法优于不纳入工具信息的传统两阶段降维方法。
原文摘要 · Abstract (English)
Traditional instrumental variable (IV) estimators face a fundamental constraint: they can only accommodate as many endogenous treatment variables as available instruments. This limitation becomes particularly challenging in settings where the treatment is presented in a high-dimensional and unstructured manner (e.g. descriptions of patient treatment pathways in a hospital). In such settings, researchers typically resort to applying unsupervised dimension reduction techniques to learn a low-dimensional treatment representation prior to implementing IV regression analysis. We show that such methods can suffer from substantial omitted variable bias due to implicit regularization in the representation learning step. We propose a novel approach to construct treatment representations by explicitly incorporating instrumental variables during the representation learning process. Our approach provides a framework for handling high-dimensional endogenous variables with limited instruments. We demonstrate both theoretically and empirically that fitting IV models on these instrument-informed representations ensures identification of directions that optimize outcome prediction. Our experiments show that our proposed methodology improves upon the conventional two-stage approaches that perform dimension reduction without incorporating instrument information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。