arXiv:2509.23570cs.LG2025-09被引 1

用可信种子和稳健传播提升因果发现准确率

Improving constraint-based discovery with robust propagation and reliable LLM priors

  • 结合置信度检验与大模型标注,筛选高可信边起点
  • 通过打乱查询识别幻觉,提升初始种子可靠性
  • 优先定向高置信边,适合真实数据中因果推断

从观测数据中学习因果结构是科学建模与决策的核心。约束方法旨在恢复因果有向无环图(DAG)中的条件独立(CI)关系。经典方法如PC先定向v-结构,再传播边方向,依赖完美CI测试和分离子集的穷尽搜索——这些假设在实践中常被违反,导致最终图中误差级联。近期研究尝试用大语言模型(LLMs)作为专家,提示节点集以确定边方向,可在假设不成立时增强定向能力。然而此类方法隐含了完美专家假设,对易产生幻觉的LLM而言不现实。本文提出MosaCD,一种因果发现方法:从CI测试与LLM标注中提取高置信度种子,通过打乱查询利用LLM的位置偏见,过滤幻觉,仅保留高置信种子。随后采用新型置信度降级传播策略,优先定向最可靠边,可集成于任意基于骨架的发现方法。在多个真实世界图上,相较于现有约束方法,MosaCD显著提升最终图的构建准确率,主要归功于初始种子可靠性的提升与稳健传播策略。

原文摘要 · Abstract (English)

Learning causal structure from observational data is central to scientific modeling and decision-making. Constraint-based methods aim to recover conditional independence (CI) relations in a causal directed acyclic graph (DAG). Classical approaches such as PC and subsequent methods orient v-structures first and then propagate edge directions from these seeds, assuming perfect CI tests and exhaustive search of separating subsets -- assumptions often violated in practice, leading to cascading errors in the final graph. Recent work has explored using large language models (LLMs) as experts, prompting sets of nodes for edge directions, and could augment edge orientation when assumptions are not met. However, such methods implicitly assume perfect experts, which is unrealistic for hallucination-prone LLMs. We propose MosaCD, a causal discovery method that propagates edges from a high-confidence set of seeds derived from both CI tests and LLM annotations. To filter hallucinations, we introduce shuffled queries that exploit LLMs' positional bias, retaining only high-confidence seeds. We then apply a novel confidence-down propagation strategy that orients the most reliable edges first, and can be integrated with any skeleton-based discovery method. Across multiple real-world graphs, MosaCD achieves higher accuracy in final graph construction than existing constraint-based methods, largely due to the improved reliability of initial seeds and robust propagation strategies.

因果发现大模型置信传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。