用大模型辅助发现因果关系,结合数据与专家知识
Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach
- 让大模型从变量名中提取结构先验,结合独立性检验结果
- 在标准基准上达到当前最优性能,优于传统方法
- 适合需要快速构建因果图的科研与工业场景
因果发现旨在从数据中揭示因果关系,通常以因果图形式表示,对干预效果预测至关重要。尽管构建严谨因果图需依赖专家知识,但已有多种统计方法基于观测数据进行推断,且具有不同形式保证。因果假设论证(Causal ABA)是一种利用符号推理确保输入约束与输出图一致的框架,能系统整合数据与专家知识。本文探索将大语言模型(LLM)作为不完美专家,从变量名称和描述中提取语义结构先验,并与条件独立性证据融合。在标准基准和语义基础的合成图上实验表明,该方法达到当前最优性能。此外,我们提出一种评估协议,以缓解评估过程中因记忆偏差带来的问题。
原文摘要 · Abstract (English)
Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the effects of interventions. While expert knowledge is required to construct principled causal graphs, many statistical methods have been proposed to leverage observational data with varying formal guarantees. Causal Assumption-based Argumentation (ABA) is a framework that uses symbolic reasoning to ensure correspondence between input constraints and output graphs, while offering a principled way to combine data and expertise. We explore the use of large language models (LLMs) as imperfect experts for Causal ABA, eliciting semantic structural priors from variable names and descriptions and integrating them with conditional-independence evidence. Experiments on standard benchmarks and semantically grounded synthetic graphs demonstrate state-of-the-art performance, and we additionally introduce an evaluation protocol to mitigate memorisation bias when assessing LLMs for causal discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。