arXiv:2511.07236cs.LG2025-11被引 7

用预训练模型提取因果关系,效果优于传统方法。

Does TabPFN Understand Causal Structures?

  • 设计可学习解码器与因果标记,从冻结嵌入中提取因果信号。
  • 中层特征包含主要因果信息,性能超过多个经典算法。
  • 为可解释的表格基础模型提供新思路,适合因果推断研究者。

因果发现对多个科学领域至关重要,但从真实数据中提取因果信息仍具挑战性。鉴于近期在真实数据上的成功,我们研究了基于Transformer的表格基础模型TabPFN是否在其内部表示中编码了因果信息。该模型在由结构因果模型生成的合成数据上预训练。我们开发了一种适配器框架,使用可学习解码器和因果标记,从TabPFN的冻结嵌入中提取因果信号,并解码为邻接矩阵以实现因果发现。评估表明,TabPFN嵌入中确实包含因果信息,在多个数据集上表现优于多种传统因果发现算法,且该信息集中于中间层。这些发现为可解释且可适应的基础模型开辟了新方向,展示了利用预训练表格模型进行因果发现的潜力。

原文摘要 · Abstract (English)

Causal discovery is fundamental for multiple scientific domains, yet extracting causal information from real world data remains a significant challenge. Given the recent success on real data, we investigate whether TabPFN, a transformer-based tabular foundation model pre-trained on synthetic datasets generated from structural causal models, encodes causal information in its internal representations. We develop an adapter framework using a learnable decoder and causal tokens that extract causal signals from TabPFN's frozen embeddings and decode them into adjacency matrices for causal discovery. Our evaluations demonstrate that TabPFN's embeddings contain causal information, outperforming several traditional causal discovery algorithms, with such causal information being concentrated in mid-range layers. These findings establish a new direction for interpretable and adaptable foundation models and demonstrate the potential for leveraging pre-trained tabular models for causal discovery.

因果发现基础模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。