arXiv:2510.09217cs.CLcs.LG2025-10ACL被引 2

无需表格数据,自动发现已知与未知因果关系。

IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular Data

  • 融合统计方法与大模型,迭代挖掘因果关系
  • 可自动发现缺失变量并扩充因果图谱
  • 适合无现成数据的科学探索场景

因果发现是科学研究的基础,但传统统计算法面临数据采集成本高、重复计算已知关系、假设不现实等问题。尽管近期基于大模型的方法能识别常见因果关系,却难以发现新关系。我们提出IRIS(迭代检索与集成系统),从一组初始变量出发,自动收集相关文献、提取变量并发现因果关系。该混合方法结合统计算法与大模型能力,既能发现已知关系,也能挖掘新关系。IRIS还包含缺失变量建议模块,可识别并引入遗漏变量以扩展因果图。本方法仅需初始变量即可实现实时因果发现,无需预先存在的数据集。

原文摘要 · Abstract (English)

Causal discovery is fundamental to scientific research, yet traditional statistical algorithms face significant challenges, including expensive data collection, redundant computation for known relations, and unrealistic assumptions. While recent LLM-based methods excel at identifying commonly known causal relations, they fail to uncover novel relations. We introduce IRIS (Iterative Retrieval and Integrated System for Real-Time Causal Discovery), a novel framework that addresses these limitations. Starting with a set of initial variables, IRIS automatically collects relevant documents, extracts variables, and uncovers causal relations. Our hybrid causal discovery method combines statistical algorithms and LLM-based methods to discover known and novel causal relations. In addition to causal discovery on initial variables, the missing variable proposal component of IRIS identifies and incorporates missing variables to expand the causal graphs. Our approach enables real-time causal discovery from only a set of initial variables without requiring pre-existing datasets.

因果发现大模型知识挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。