文本因果推断中存在治疗泄漏问题,本文提出四种去噪方法有效缓解偏倚。
Detecting and Mitigating Treatment Leakage in Text-Based Causal Inference: Distillation and Sensitivity Analysis
- 通过四类文本提炼技术去除预测治疗状态的内容
- 适度提炼可平衡偏倚降低与混杂因子保留
- 适用于文本作为混杂因子的因果分析场景
基于文本的因果推断越来越多地将文本数据用作未观测混杂因子的代理,但这一方法引入了一种此前未被充分理论化的偏差来源:治疗泄漏。治疗泄漏指本用于捕捉混杂信息的文本中也包含预测治疗状态的信号,从而在因果估计中引发治疗后偏倚。关键的是,即使文档在治疗分配之前生成,作者也可能使用未来引用语言提前预示后续干预,导致该问题出现。尽管该问题日益受到关注,但目前尚无系统方法用于识别和缓解文本作为混杂因子时的治疗泄漏。本文通过三项贡献填补这一空白:第一,提供治疗泄漏的形式化统计与集合论定义,明确偏倚发生的时间与原因;第二,提出四种文本提炼方法——基于相似性的段落剔除、远监督分类、显著特征移除和迭代零空间投影,旨在消除治疗预测内容的同时保留混杂信息;第三,通过合成文本模拟和一个实证研究(分析国际货币基金组织结构调整计划与儿童死亡率的关系)验证这些方法。结果表明,适度提炼可在降低偏倚与保持混杂因子信息之间取得最佳平衡,而过度严格的处理会损害估计精度。
原文摘要 · Abstract (English)
Text-based causal inference increasingly employs textual data as proxies for unobserved confounders, yet this approach introduces a previously undertheorized source of bias: treatment leakage. Treatment leakage occurs when text intended to capture confounding information also contains signals predictive of treatment status, thereby inducing post-treatment bias in causal estimates. Critically, this problem can arise even when documents precede treatment assignment, as authors may employ future-referencing language that anticipates subsequent interventions. Despite growing recognition of this issue, no systematic methods exist for identifying and mitigating treatment leakage in text-as-confounder applications. This paper addresses this gap through three contributions. First, we provide formal statistical and set-theoretic definitions of treatment leakage that clarify when and why bias occurs. Second, we propose four text distillation methods -- similarity-based passage removal, distant supervision classification, salient feature removal, and iterative nullspace projection -- designed to eliminate treatment-predictive content while preserving confounder information. Third, we validate these methods through simulations using synthetic text and an empirical application examining International Monetary Fund structural adjustment programs and child mortality. Our findings indicate that moderate distillation optimally balances bias reduction against confounder retention, whereas overly stringent approaches degrade estimate precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。