arXiv:2601.01511cs.AI2026-01

用文本嵌入消除因果估计中的隐藏偏误,深度学习更有效。

Reading Between the Lines: Deconfounding Causal Estimates using Text Embeddings and Deep Learning

  • 用文本嵌入捕捉结构化数据遗漏的混杂因子
  • 传统树模型偏差高达+24%,深度学习降至-0.86%
  • 适合处理高维自然语言数据的因果推断任务

在观测研究中,未观测到的混杂因子常导致因果效应估计出现选择偏误。尽管传统计量方法在混杂因子与结构化协变量正交时表现不佳,但高维非结构化文本往往包含这些潜在变量的丰富代理信息。本文提出一种神经网络增强的双重机器学习(DML)框架,利用文本嵌入实现因果识别。通过严格的合成基准测试,我们证明文本嵌入能捕获结构化表格数据中缺失的关键混杂信息。然而,标准树基DML估计器因无法建模嵌入流形的连续拓扑,仍存在显著偏误(+24%)。相比之下,优化架构的深度学习方法将偏误降至-0.86%,有效恢复真实因果参数。结果表明,在基于高维自然语言数据进行条件控制时,深度学习架构对满足无混杂假设至关重要。

原文摘要 · Abstract (English)

Estimating causal treatment effects in observational settings is frequently compromised by selection bias arising from unobserved confounders. While traditional econometric methods struggle when these confounders are orthogonal to structured covariates, high-dimensional unstructured text often contains rich proxies for these latent variables. This study proposes a Neural Network-Enhanced Double Machine Learning (DML) framework designed to leverage text embeddings for causal identification. Using a rigorous synthetic benchmark, we demonstrate that unstructured text embeddings capture critical confounding information that is absent from structured tabular data. However, we show that standard tree-based DML estimators retain substantial bias (+24%) due to their inability to model the continuous topology of embedding manifolds. In contrast, our deep learning approach reduces bias to -0.86% with optimized architectures, effectively recovering the ground-truth causal parameter. These findings suggest that deep learning architectures are essential for satisfying the unconfoundedness assumption when conditioning on high-dimensional natural language data

因果推断文本嵌入深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。