用多智能体协作挖掘潜变量,让因果发现更可追溯、更可信。
Traceable Latent Variable Discovery Based on Multi-Agent Collaboration
- 通过多大模型协作建模,将潜变量推断转化为不完全信息博弈求贝叶斯均衡。
- 在五组真实医疗与基准数据上,准确率、类别准确率和证据置信度平均提升超26%。
- 适合需要可解释因果推理的医疗、社会科学等高可信场景研究者。
揭示现实世界中的潜在因果机制对科技发展至关重要。尽管近年进展显著,但高质量数据缺乏、传统因果发现算法(TCDA)依赖无潜混杂假设,且常忽略潜变量语义,长期制约其应用。为此,我们提出新框架TLVD,融合大语言模型(LLM)的元数据推理能力与TCDA的数据驱动建模能力,以推断潜变量及其语义。首先,采用数据驱动方法构建含潜变量的因果图;其次,利用多LLM协作进行潜变量推断,将其建模为不完全信息博弈,求解贝叶斯纳什均衡以获得可能的具体潜变量;最后,通过LLM进行跨多个真实网络数据源的证据探索,验证并确保推断结果的可追溯性。我们在三个脱敏患者数据集及两个基准数据集上全面评估TLVD。实验结果表明其有效且可靠,五组数据上平均提升32.67%(Acc)、62.21%(CAcc)和26.72%(ECit)。
原文摘要 · Abstract (English)
Revealing the underlying causal mechanisms in the real world is crucial for scientific and technological progress. Despite notable advances in recent decades, the lack of high-quality data and the reliance of traditional causal discovery algorithms (TCDA) on the assumption of no latent confounders, as well as their tendency to overlook the precise semantics of latent variables, have long been major obstacles to the broader application of causal discovery. To address this issue, we propose a novel causal modeling framework, TLVD, which integrates the metadata-based reasoning capabilities of large language models (LLMs) with the data-driven modeling capabilities of TCDA for inferring latent variables and their semantics. Specifically, we first employ a data-driven approach to construct a causal graph that incorporates latent variables. Then, we employ multi-LLM collaboration for latent variable inference, modeling this process as a game with incomplete information and seeking its Bayesian Nash Equilibrium (BNE) to infer the possible specific latent variables. Finally, to validate the inferred latent variables across multiple real-world web-based data sources, we leverage LLMs for evidence exploration to ensure traceability. We comprehensively evaluate TLVD on three de-identified real patient datasets provided by a hospital and two benchmark datasets. Extensive experimental results confirm the effectiveness and reliability of TLVD, with average improvements of 32.67% in Acc, 62.21% in CAcc, and 26.72% in ECit across the five datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。