arXiv:2506.18285cs.LGcs.AI2025-06被引 3

用基础模型方法高效学习大规模因果图,提升小样本下的准确性与泛化性。

Learning Causal Graphs at Scale: A Foundation Model Approach

  • 基于注意力机制构建可多任务学习的线性结构方程模型
  • 在合成数据集上显著提升因果图识别准确率和零样本推理效率
  • 适合需要快速、泛化性强因果推断的研究者使用

由于其人类可解释性和不变性特性,有向无环图(DAG)已成为人工智能多个领域的基础工具,推动了显著进展。然而,由于计算成本超指数增长及可识别性问题,尤其是在小样本情况下,DAG学习仍极具挑战。为此,本文利用线性变换器的最新成果,提出一种用于跨任务发现多个有序一致因果图的基础模型方法。我们设计了注意力因果图(ADAG),一种基于注意力机制的新架构,用于学习多个线性结构方程模型(SEMs)。ADAG通过非线性注意力核将观测数据映射到图结构与参数,实现对底层线性SEMs的高效多任务估计。通过将多任务学习过程建模为连续优化问题,预训练的ADAG模型捕捉共同结构特性作为共享低维先验,从而缓解下游小样本情形下DAG学习的病态性。我们在基准合成数据集上评估该方法,结果表明ADAG在因果图学习精度和零样本推理效率方面均有显著提升。据我们所知,这是首个专为因果图学习设计的实用预训练基础模型,标志着向更高效、更通用的因果发现应用迈进一步。

原文摘要 · Abstract (English)

Due to its human-interpretability and invariance properties, Directed Acyclic Graph (DAG) has been a foundational tool across various areas of AI research, leading to significant advancements. However, DAG learning remains highly challenging, due to its super-exponential growth in computational cost and identifiability issues, particularly in small-sample regimes. To address these two challenges, in this work we leverage the recent success of linear transformers and develop a foundation model approach for discovering multiple order-consistent DAGs across tasks. In particular, we propose Attention-DAG (ADAG), a novel attention-mechanism-based architecture for learning multiple linear Structural Equation Models (SEMs). ADAG learns the mapping from observed data to both graph structure and parameters via a nonlinear attention-based kernel, enabling efficient multi-task estimation of the underlying linear SEMs. By formulating the learning process across multiple tasks as a continuous optimization problem, the pre-trained ADAG model captures the common structural properties as a shared low-dimensional prior, thereby reducing the ill-posedness of downstream DAG learning tasks in small-sample regimes. We evaluate our proposed approach on benchmark synthetic datasets and find that ADAG achieves substantial improvements in both DAG learning accuracy and zero-shot inference efficiency. To the best of our knowledge, this is the first practical approach for pre-training a foundation model specifically designed for DAG learning, representing a step toward more efficient and generalizable down-stream applications in causal discovery.

因果推断基础模型图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。