arXiv:2602.20333cs.AI2026-02

用语义+统计双阶段发现因果关系,提升真实数据上的准确率。

DMCD: Semantic-Statistical Framework for Causal Discovery

  • 先用大模型分析变量元数据生成初步因果图,再用统计检验修正
  • 在三个真实数据集上召回率和F1分数显著优于现有方法
  • 适合需要高可信度因果推断的工业、环境等场景

我们提出DMCD(DataMap因果发现)框架,采用两阶段方法:第一阶段利用大语言模型基于变量元数据生成稀疏的初始因果图,作为对可能因果结构的语义先验;第二阶段通过条件独立性检验对初稿进行审计与优化,检测到的偏差用于指导针对性的边修改。我们在涵盖工业工程、环境监测和信息系统分析的三个元数据丰富的实际数据集上评估该方法。结果表明,DMCD在多个基准上表现优异,尤其在召回率和F1分数上取得显著提升。探查与消融实验显示,性能提升源于对元数据的语义推理,而非对基准图的记忆。总体而言,结合语义先验与严谨统计验证可实现高效且实用的因果结构学习。

原文摘要 · Abstract (English)

We present DMCD (DataMap Causal Discovery), a two-phase causal discovery framework that integrates LLM-based semantic drafting from variable metadata with statistical validation on observational data. In Phase I, a large language model proposes a sparse draft DAG, serving as a semantically informed prior over the space of possible causal structures. In Phase II, this draft is audited and refined via conditional independence testing, with detected discrepancies guiding targeted edge revisions. We evaluate our approach on three metadata-rich real-world benchmarks spanning industrial engineering, environmental monitoring, and IT systems analysis. Across these datasets, DMCD achieves competitive or leading performance against diverse causal discovery baselines, with particularly large gains in recall and F1 score. Probing and ablation experiments suggest that these improvements arise from semantic reasoning over metadata rather than memorization of benchmark graphs. Overall, our results demonstrate that combining semantic priors with principled statistical verification yields a high-performing and practically effective approach to causal structure learning.

因果发现大模型数据分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。