首次系统评估因果发现基准图的可靠性,发现主流基准存在知识不一致问题。
Consistency evaluation of benchmarks used for causal discovery
- 用LLM自动检索论文并比对基准图与领域研究的一致性
- 分析11个基准,处理38,081篇文献,发现一致性差异显著
- 对基于大模型的因果发现方法评估有重要影响
在图形因果模型中,因果发现旨在根据数值数据和文本领域知识构建因果图。然而,由于领域研究进展导致基准因果图常包含错位知识,因果发现方法的评估仍面临挑战。这一问题尤其影响基于大语言模型(LLM)的因果发现方法,因其对新文献发现敏感。本文首次系统研究基准因果图的质量。我们设计了一个自动化流程,从科学数据库中检索相关论文,并通过提示LLM检查基准图与领域论文之间的一致性。我们评估了11个流行的现实世界基准,共处理38,081篇领域论文。结果表明,主流基准在与领域研究的一致性上差异显著,对因果发现研究具有明确启示。
原文摘要 · Abstract (English)
In graphical causal model, causal discovery aims to construct a causal graph based on numerical data and domain knowledge in plain text. However, the evaluation of causal discovery methods remains a challenge in the area as the progress of domain researches often makes benchmark causal graphs contain mis-aligned knowledge. This problem especially affects the evaluation of large language model (LLM) based causal discovery methods as they are sensitive to the new discoveries in the literature. This work is the first to systematically study the quality of benchmark causal graphs. Specifically, we design a pipeline that automatically retrieves relevant research papers from scientific databases, and prompts LLMs to check the consistency between the benchmark causal graphs and domain research papers. We evaluate 11 popular real-world benchmarks, for which our pipeline in total proceeds 38,081 domain papers. Our results show that popular benchmarks vary significantly in their consistency with domain research, with clear implications for causal discovery research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。