用多智能体大模型自动发现因果推断中的有效工具变量
IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery
- 设计多智能体系统,让不同角色的LLM协作提出、批判和优化工具变量
- 在无真实答案的情况下通过统计测试验证结果一致性,提升可靠性
- 适用于经济学、流行病学等需要因果推断的研究者
当内生变量与结果变量存在混杂偏差时,工具变量(IV)可用于分离其因果效应。识别有效的工具变量需跨学科知识、创造力和上下文理解,是一项复杂任务。本文探讨大语言模型(LLMs)是否能辅助该任务。采用两阶段评估框架:第一阶段测试LLMs能否从文献中复现已知有效工具变量,评估其推理能力;第二阶段检验其能否识别并规避已被实证或理论否定的工具变量。基于上述结果,提出IV Co-Scientist——一个针对给定处理-结果对生成、批评与优化工具变量的多智能体系统。同时引入统计检验方法,在缺乏真实答案时判断结果的一致性。实验表明,LLMs具备从大规模观察数据中发现有效工具变量的潜力。
原文摘要 · Abstract (English)
In the presence of confounding between an endogenous variable and the outcome, instrumental variables (IVs) are used to isolate the causal effect of the endogenous variable. Identifying valid instruments requires interdisciplinary knowledge, creativity, and contextual understanding, making it a non-trivial task. In this paper, we investigate whether large language models (LLMs) can aid in this task. We perform a two-stage evaluation framework. First, we test whether LLMs can recover well-established instruments from the literature, assessing their ability to replicate standard reasoning. Second, we evaluate whether LLMs can identify and avoid instruments that have been empirically or theoretically discredited. Building on these results, we introduce IV Co-Scientist, a multi-agent system that proposes, critiques, and refines IVs for a given treatment-outcome pair. We also introduce a statistical test to contextualize consistency in the absence of ground truth. Our results show the potential of LLMs to discover valid instrumental variables from a large observational database.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。