arXiv:2410.21141cs.LGstat.ML2024-10被引 7

用大模型初始化因果发现,提升可解释性与准确性。

LLM-initialized Differentiable Causal Discovery

  • 用大模型生成初始因果图,优化最大似然目标
  • 在基准数据集上准确率优于现有方法
  • 适合需要可解释因果推理的研究者

发现随机变量间的因果关系是众多科学领域中的重要但具挑战性的问题。可微分因果发现(DCD)方法能从观测数据中揭示因果关系,但常缺乏可解释性,且难以融入领域先验知识。相比之下,基于大语言模型(LLMs)的因果发现方法虽能提供有用先验,却在形式化因果推理方面存在不足。本文提出LLM-DCD,利用大模型初始化DCD方法的最大似然目标函数优化过程,从而将强先验知识引入发现流程。为此,我们设计目标函数仅以显式定义的因果图邻接矩阵为变分参数,直接优化该矩阵使方法更具可解释性。实验表明,本方法在关键基准数据集上的准确率高于现有最优方法,并实证验证了初始化质量直接影响最终结果。LLM-DCD为传统因果发现方法(如DCD)未来受益于大模型因果推理能力提升开辟了新路径。

原文摘要 · Abstract (English)

The discovery of causal relationships between random variables is an important yet challenging problem that has applications across many scientific domains. Differentiable causal discovery (DCD) methods are effective in uncovering causal relationships from observational data; however, these approaches often suffer from limited interpretability and face challenges in incorporating domain-specific prior knowledge. In contrast, Large Language Models (LLMs)-based causal discovery approaches have recently been shown capable of providing useful priors for causal discovery but struggle with formal causal reasoning. In this paper, we propose LLM-DCD, which uses an LLM to initialize the optimization of the maximum likelihood objective function of DCD approaches, thereby incorporating strong priors into the discovery method. To achieve this initialization, we design our objective function to depend on an explicitly defined adjacency matrix of the causal graph as its only variational parameter. Directly optimizing the explicitly defined adjacency matrix provides a more interpretable approach to causal discovery. Additionally, we demonstrate higher accuracy on key benchmarking datasets of our approach compared to state-of-the-art alternatives, and provide empirical evidence that the quality of the initialization directly impacts the quality of the final output of our DCD approach. LLM-DCD opens up new opportunities for traditional causal discovery methods like DCD to benefit from future improvements in the causal reasoning capabilities of LLMs.

因果发现大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。