arXiv:2605.31156cs.LG2026-05被引 1

通过多样化因果环境预训练,提升表格数据因果发现的泛化能力。

TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery

论文配图:TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery
图 1 · 摘自论文原文
  • 基于动态任务构造,融合多种因果结构与干预模式进行预训练。
  • 在合成数据上性能超越多种基线方法,尤其在干预数据下表现优异。
  • 适合需要跨场景可靠因果推理的研究者与工业应用者。

因果发现旨在从观测与干预数据中恢复有向因果关系,为机制理解与可靠决策提供基础。因果发现基础模型(CDFM)试图通过单次前向传播将数据映射为因果图,避免逐数据集测试与优化。然而现有CDFM仍受限,常无法稳定匹配经典方法,我们发现关键瓶颈在于因果预训练任务的设计。为此,提出TabCausal,一种基于数据驱动的CDFM,其在多样化的图先验、结构机制、噪声模型、维度、样本量及干预策略上进行广泛因果预训练。采用动态任务构建策略,将这些因果环境组合成多样的发现任务,从而实现从观测与混合干预数据中更具迁移性的结构学习。在大规模合成基准上,TabCausal的宏平均性能优于多种因果发现基线。为进一步连接抽象合成生成器与真实因果推理场景,引入协议引导且经大语言模型审核的语义因果环境基准,其中领域相关结构因果模型生成可解释的观测与干预数据,用于分布外分析。在合成与语义环境中,TabCausal均展现出稳健的结构恢复能力,尤其在干预证据下表现突出,凸显广泛因果预训练对可迁移的摊销式因果发现的重要性。

原文摘要 · Abstract (English)

Causal discovery aims to recover directed causal relations from observational and interventional data, providing a basis for mechanistic understanding and reliable decision-making. Causal discovery foundation models (CDFMs) seek to amortize this problem by mapping a dataset directly to a causal graph in a single forward pass, avoiding per-dataset testing, search, or optimization. However, existing CDFMs remain limited, often failing to consistently match strong classical methods, and we find that a key bottleneck is how causal pretraining tasks are constructed. Based on this observation, we propose TabCausal, a data-driven CDFM trained with broad causal pretraining over diverse graph priors, structural mechanisms, noise models, dimensions, sample sizes, and intervention regimes. A dynamic task construction strategy composes these causal environments into varied discovery tasks, enabling more transferable structural learning from observational and mixed-interventional data. On large-scale synthetic benchmarks, TabCausal achieves better macro-averaged performance than a diverse set of causal discovery baselines. To further bridge abstract synthetic generators and realistic causal reasoning scenarios, we introduce a protocol-guided and LLM-audited semantic causal environment benchmark, where domain-grounded SCMs generate interpretable observational and interventional datasets for out-of-distribution analysis. Across both synthetic and semantic environments, TabCausal demonstrates robust structure recovery, especially under interventional evidence, highlighting broad causal pretraining as a key ingredient for transferable amortized causal discovery.

因果发现预训练表格数据结构学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。