arXiv:2607.26309stat.MEcs.CL2026-07

文本中治疗词会干扰因果推断,提出掩码法去除干扰信号

The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text

论文配图:The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text
图 1 · 摘自论文原文
  • 用掩码去除文本中治疗相关词汇,避免表示学习时泄露治疗信息
  • 实验显示掩码后重叠性诊断改善,处理效应估计更稳定且偏差更低
  • 适用于基于文本的因果推断,尤其当治疗由关键词定义时

从观察性文本中估计语言属性的因果效应极具挑战,因为同一文档可能同时包含感兴趣处理和需调整的非处理文本特征。现有方法通常从全文学习表示以捕捉潜在混杂,但当处理状态本身通过文本中的词语编码时,这些表示会直接包含处理信号。这导致‘混杂陷阱’:更丰富的表示反而使处理组与对照组可分,引发重叠性违反,即使原始因果问题满足重叠条件。本文研究通过词典或其他词法信息编码的隐含文本处理,提出基于掩码的调整表示方法,在表示学习前移除此类词法处理信号。我们形式化了表示引起的重叠失败,证明删除掩码能保持词袋/主题模型表示的重叠性,并将替换掩码视为大语言模型的自然松弛——隐藏治疗定义词但保留词序与上下文。在模拟实验中,掩码方法显著改善重叠性诊断,稳定处理效应估计,降低偏差,优于从未掩码文本学习的方法。

原文摘要 · Abstract (English)

Estimating causal effects of linguistic properties from observational text is difficult because the same document can contain both the treatment of interest and the non-treatment textual attributes needed for adjustment. Existing approaches often learn representations from the full text to capture latent confounding, but when treatment status is itself encoded by words in the text, these representations can directly encode treatment. This creates a confounder trap: richer representations can make treated and control documents separable, inducing overlap violations even when the underlying causal problem satisfies overlap. We study latent text treatments that are encoded through lexicons or other treatment-defining lexical information, and propose masking-based adjustment representations that remove this lexical treatment signal before representation learning. We formalize representation-induced overlap failure, prove that deletion masking preserves overlap for bag-of-words/topic-model representations, and characterize replacement masking as a natural relaxation for large language models that hides treatment-defining tokens while preserving word order and context. Across simulations, masking improves overlap diagnostics, stabilizes treatment effect estimates, and reduces bias relative to adjustment methods that learn from the unmasked text.

因果推断文本分析表示学习掩码机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。