大模型不能发现因果关系,只能辅助搜索过程。
LLM Cannot Discover Causality, and Should Be Restricted to Non-Decisional Support in Causal Discovery
- 限制大模型仅作非决策性辅助,不参与因果判断。
- 实验证明,隔离大模型可加速收敛并超越传统方法。
- 适合关注因果推断可靠性的研究者参考。
本文重新审视大模型在因果发现中的作用,主张其不应直接参与确定因果关系。我们证明大模型的自回归、相关性驱动建模缺乏因果推理的理论基础,在作为因果发现算法先验时引入不可靠性。通过实证研究,揭示现有基于大模型的方法存在局限,刻意提示工程(如注入真实知识)可能夸大其性能,解释了当前文献中普遍呈现的乐观结果。基于此,我们严格限定大模型角色为非决策性辅助:不能决定因果关系的存在或方向,但可协助因果图搜索(如大模型启发式搜索)。跨多种设置的实验表明,通过将大模型严格排除在因果决策之外,大模型引导的启发式搜索能加速收敛,并优于传统及现有大模型方法。最后呼吁学界转向开发符合因果发现核心原则的专用模型与训练方法。
原文摘要 · Abstract (English)
This paper critically re-evaluates LLMs' role in causal discovery and argues against their direct involvement in determining causal relationships. We demonstrate that LLMs' autoregressive, correlation-driven modeling inherently lacks the theoretical grounding for causal reasoning and introduces unreliability when used as priors in causal discovery algorithms. Through empirical studies, we expose the limitations of existing LLM-based methods and reveal that deliberate prompt engineering (e.g., injecting ground-truth knowledge) could overstate their performance, helping to explain the consistently favorable results reported in much of the current literature. Based on these findings, we strictly confined LLMs' role to a non-decisional auxiliary capacity: LLMs should not participate in determining the existence or directionality of causal relationships, but can assist the search process for causal graphs (e.g., LLM-based heuristic search). Experiments across various settings confirm that, by strictly isolating LLMs from causal decision-making, LLM-guided heuristic search can accelerate the convergence and outperform both traditional and LLM-based methods in causal structure learning. We conclude with a call for the community to shift focus from naively applying LLMs to developing specialized models and training method that respect the core principles of causal discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。