提出科学基准与混合方法,让大模型真正具备因果推理能力
Realizing LLMs' Causal Potential Requires Science-Grounded, Novel Benchmarks
- 用训练后发表的科研论文构建新基准,避免数据泄露
- 大模型在真实科学图谱上表现远差于旧基准,暴露记忆依赖
- 结合大模型先验与统计算法,显著提升因果发现准确率
近期声称大语言模型(LLMs)在因果发现上表现优异的说法受到质疑:许多评估所用基准可能已出现在预训练语料中。这导致看似成功的性能,实则反映模型对观测数据的忽视,而非真正的因果推理。我们提出两个关键转变:(P.1) 建立基于近期科学文献的稳健评估协议,防止数据泄露;(P.2) 设计融合大模型知识与数据驱动统计的混合方法。为此,我们建议在大模型训练截止时间后的最新科研论文中提取因果图,确保新颖性与无记忆性,同时捕捉已有与新发现的关系。相比在BNLearn等基准上接近完美的表现,大模型在新构建图谱上的性能显著下降。此外,将大模型预测作为经典PC算法的先验信息,可显著优于纯大模型或纯统计方法。我们呼吁社区采用科学支撑、抗泄露的基准,并发展适用于真实科学探究的混合因果发现方法。
原文摘要 · Abstract (English)
Recent claims of strong performance by Large Language Models (LLMs) on causal discovery are undermined by a key flaw: many evaluations rely on benchmarks likely included in pretraining corpora. Thus, apparent success suggests that LLM-only methods, which ignore observational data, outperform classical statistical approaches. We challenge this narrative by asking: Do LLMs truly reason about causal structure, and how can we measure it without memorization concerns? Can they be trusted for real-world scientific discovery? We argue that realizing LLMs' potential for causal analysis requires two shifts: (P.1) developing robust evaluation protocols based on recent scientific studies to guard against dataset leakage, and (P.2) designing hybrid methods that combine LLM-derived knowledge with data-driven statistics. To address P.1, we encourage evaluating discovery methods on novel, real-world scientific studies. We outline a practical recipe for extracting causal graphs from recent publications released after an LLM's training cutoff, ensuring relevance and preventing memorization while capturing both established and novel relations. Compared to benchmarks like BNLearn, where LLMs achieve near-perfect accuracy, they perform far worse on our curated graphs, underscoring the need for statistical grounding. Supporting P.2, we show that using LLM predictions as priors for the classical PC algorithm significantly improves accuracy over both LLM-only and purely statistical methods. We call on the community to adopt science-grounded, leakage-resistant benchmarks and invest in hybrid causal discovery methods suited to real-world inquiry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。