arXiv:2512.05373cs.LGcs.AI2025-12被引 1

通过筛选关键文本词元,提升因果效应估计的准确性与稳定性。

Text Rationalization for Robust Causal Effect Estimation

  • 用残差独立性诊断筛选必要词元,保留混淆因子信息
  • 在真实医疗数据上降低倾向得分极端值,减少估计方差
  • 适合需要可解释因果推断的医疗、社会科学领域

自然语言处理的进步推动了文本数据在因果推断中的应用,尤其用于调整治疗效应估计中的混杂因素。高维文本虽蕴含丰富上下文信息,却也带来因果识别挑战。观测水平上常违反正性假设(positivity assumption),即不同混杂值间治疗重叠不足,大量冗余或虚假文本特征会膨胀维度,导致极端倾向得分、不稳定的权重和效应估计方差增大。本文提出混淆感知词元合理化框架(CATR),利用残差独立性诊断筛选稀疏但必要的词元,保留足够实现无混杂性的信息。通过剔除无关文本并保留关键信号,CATR缓解了观测水平上的正性违背问题,稳定了下游因果效应估计器。在合成数据和基于MIMIC-III数据库的真实世界研究中,CATR比现有基线方法获得更准确、更稳定且可解释的因果效应估计。

原文摘要 · Abstract (English)

Recent advances in natural language processing have enabled the increasing use of text data in causal inference, particularly for adjusting confounding factors in treatment effect estimation. Although high-dimensional text can encode rich contextual information, it also poses unique challenges for causal identification and estimation. In particular, the positivity assumption, which requires sufficient treatment overlap across confounder values, is often violated at the observational level, when massive text is represented in feature spaces. Redundant or spurious textual features inflate dimensionality, producing extreme propensity scores, unstable weights, and inflated variance in effect estimates. We address these challenges with Confounding-Aware Token Rationalization (CATR), a framework that selects a sparse necessary subset of tokens using a residual-independence diagnostic designed to preserve confounding information sufficient for unconfoundedness. By discarding irrelevant texts while retaining key signals, CATR mitigates observational-level positivity violations and stabilizes downstream causal effect estimators. Experiments on synthetic data and a real-world study using the MIMIC-III database demonstrate that CATR yields more accurate, stable, and interpretable causal effect estimates than existing baselines.

因果推断文本分析医疗数据词元筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。