arXiv:2506.16234cs.LG2025-06

用语言模型辅助因果发现,能处理数据分批到达和专家知识不足的问题。

Sequential Causal Discovery with Noisy Language Model Priors

  • 将语言模型的先验知识与分批观测数据动态结合,修正两类偏差。
  • 在多个数据集上结构准确率超越现有方法,对语言模型噪声有鲁棒性。
  • 适合缺乏完整数据或专家资源的研究者使用,尤其适用于动态数据场景。

从观测数据中进行因果发现通常假设数据完整且存在完美领域专家。然而现实中,数据常以批次形式到达,受采样偏差影响,且专家知识稀缺。语言模型(LM)可作为专家知识的替代,但存在幻觉、不一致和偏见问题。本文提出一种混合框架,通过自适应融合序列批次数据与语言模型生成的有噪声先验知识,同时考虑数据和语言模型带来的双重偏差。我们提出将有向无环图(DAG)表示转换为部分祖先图(PAG),在统一框架内处理不确定性,实现全局语言模型知识与局部观测数据的对齐。为指导语言模型交互,采用序列优化策略,自适应查询最具有信息量的因果边。在多种数据集和语言模型上,本方法在结构准确性方面优于先前工作,并扩展至参数估计任务,展现出对语言模型噪声的鲁棒性。

原文摘要 · Abstract (English)

Causal discovery from observational data typically assumes access to complete data and availability of perfect domain experts. In practice, data often arrive in batches, are subject to sampling bias, and expert knowledge is scarce. Language Models (LMs) offer a surrogate for expert knowledge but suffer from hallucinations, inconsistencies, and bias. We present a hybrid framework that bridges these gaps by adaptively integrating sequential batch data with LM-derived noisy, expert knowledge while accounting for both data-induced and LM-induced biases. We propose a representation shift from Directed Acyclic Graph (DAG) to Partial Ancestral Graph (PAG), that accommodates ambiguities within a coherent framework, allowing grounding the global LM knowledge in local observational data. To guide LM interactions, we use a sequential optimization scheme that adaptively queries the most informative edges. Across varied datasets and LMs, we outperform prior work in structural accuracy and extend to parameter estimation, showing robustness to LM noise.

因果发现语言模型数据分批鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。