用大模型逐步推断隐藏混杂因素,提升因果推断准确性。
Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models
- 利用大模型语义推理与世界知识,迭代生成并验证隐藏混杂变量。
- 在多个数据集上显著改善处理效应估计,降低偏倚。
- 适合从事因果推断、医疗分析的研究者使用。
隐藏混杂是基于观察数据估计处理效应的核心挑战,未观测变量可能导致因果估计偏差。尽管近期研究探索了大语言模型(LLM)在因果推断中的应用,但多数方法仍依赖无混杂假设。本文首次提出使用大模型缓解隐藏混杂,引入ProCI(渐进式混杂因子插补)框架,通过大模型的语义推理能力与嵌入的世界知识,从结构化和非结构化输入中发现潜在混杂因子,并迭代生成、插补与验证。为增强鲁棒性,ProCI采用分布推理策略而非直接值插补,防止输出坍缩。大量实验表明,ProCI能有效发现有意义的混杂因子,并在多种数据集和大模型上显著提升处理效应估计性能。
原文摘要 · Abstract (English)
Hidden confounding remains a central challenge in estimating treatment effects from observational data, as unobserved variables can lead to biased causal estimates. While recent work has explored the use of large language models (LLMs) for causal inference, most approaches still rely on the unconfoundedness assumption. In this paper, we make the first attempt to mitigate hidden confounding using LLMs. We propose ProCI (Progressive Confounder Imputation), a framework that elicits the semantic and world knowledge of LLMs to iteratively generate, impute, and validate hidden confounders. ProCI leverages two key capabilities of LLMs: their strong semantic reasoning ability, which enables the discovery of plausible confounders from both structured and unstructured inputs, and their embedded world knowledge, which supports counterfactual reasoning under latent confounding. To improve robustness, ProCI adopts a distributional reasoning strategy instead of direct value imputation to prevent the collapsed outputs. Extensive experiments demonstrate that ProCI uncovers meaningful confounders and significantly improves treatment effect estimation across various datasets and LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。