用大模型生成混杂因素并自动验证迭代,提升因果推断准确性
VIGOR+: Iterative Confounder Generation and Validation via LLM-CEVAE Feedback Loop
- 通过大模型生成混杂因素,再用CEVAE模型验证并反馈改进
- 迭代优化后混杂因素在统计上更有效,信息增益提升显著
- 适合做因果推断、医疗数据分析的研究者使用
隐藏混杂仍是观测数据中因果推断的核心挑战。现有方法利用大语言模型(LLM)基于领域知识生成合理的隐藏混杂因素,但存在一个关键问题:生成的混杂因素常具语义合理性,却缺乏统计效用。我们提出VIGOR+(变分信息增益用于迭代混杂因子优化),一种将基于大模型的混杂生成与基于CEVAE的统计验证闭环结合的新框架。与以往将生成和验证视为独立阶段的方法不同,VIGOR+建立了一个迭代反馈机制:从CEVAE获得的验证信号(包括信息增益、潜在变量一致性度量及诊断消息)被转化为自然语言反馈,指导后续大模型生成。该过程持续进行直至满足收敛条件。我们形式化了反馈机制,证明在弱假设下具备收敛性,并提供了完整的算法框架。
原文摘要 · Abstract (English)
Hidden confounding remains a fundamental challenge in causal inference from observational data. Recent advances leverage Large Language Models (LLMs) to generate plausible hidden confounders based on domain knowledge, yet a critical gap exists: LLM-generated confounders often exhibit semantic plausibility without statistical utility. We propose VIGOR+ (Variational Information Gain for iterative cOnfounder Refinement), a novel framework that closes the loop between LLM-based confounder generation and CEVAE-based statistical validation. Unlike prior approaches that treat generation and validation as separate stages, VIGOR+ establishes an iterative feedback mechanism: validation signals from CEVAE (including information gain, latent consistency metrics, and diagnostic messages) are transformed into natural language feedback that guides subsequent LLM generation rounds. This iterative refinement continues until convergence criteria are met. We formalize the feedback mechanism, prove convergence properties under mild assumptions, and provide a complete algorithmic framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。