提出新模型TabFORGE,高效生成结构保真的表格数据。
Tabular Foundation Model for Generative Modelling

- 用预训练因果编码器统一建模异构表格的潜在结构
- 两阶段设计:扩散变换器+去噪解码器,提升生成质量
- 在45个真实数据集上超越22种方法,结构保真度高
生成建模是对基础模型的严苛考验,要求对特定数据模态具备鲁棒且整体的表征学习能力,而非仅优化监督预测目标。尽管近期表格基础模型在预测建模方面取得显著进展,但生成式表格基础模型仍研究不足。现有表格生成模型尚未一致超越特定数据集的强生成器,关键原因在于其与异构表格数据特有的因果结构先验不匹配。本文提出新型表格基础模型TabFORGE,基于预训练的表征学习框架,利用统一潜在空间中隐含的因果信息。通过两阶段设计:先预训练基于得分的扩散变压器,再使用去噪后的潜在嵌入预训练对齐解码器,有效缓解训练与推理间潜在分布偏移。我们在45个真实世界数据集上全面评估了TabFORGE,对比22种基准方法。结果表明,TabFORGE能有效学习通用表格表征,实现高质量合成数据的高效生成,尤其在结构保真度方面表现突出。
原文摘要 · Abstract (English)
Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, rather than optimisation for a supervised prediction target alone. While recent work on tabular foundation models has achieved remarkable progress in predictive modelling, generative tabular foundation models remain underexplored. Existing tabular foundation generators, in particular, have not yet consistently matched strong dataset-specific generators in synthetic data quality. A key reason is their misalignment with the distinctive causal structural prior of heterogeneous tabular data. In this paper, we address this gap by introducing a novel tabular foundation model, \textbf{TabFORGE}, built on pretrained \textbf{Tab}ular \textbf{FO}undational \textbf{R}epresentations for \textbf{GE}neration. TabFORGE is designed to utilise the implicitly learned causal information underlying diverse tabular datasets in a unified latent space induced by a pretrained causality-aware feature encoder. It further decouples latent modelling from decoding through a two-stage design: we first pretrain a score-based diffusion transformer, and then pretrain a denoising-aligned decoder using the denoised latent embeddings. This design elegantly mitigates the distribution shifts in latent embeddings that typically arise between training and inference. We evaluate TabFORGE comprehensively against 22 benchmark methods on 45 real-world datasets. Our results show that TabFORGE effectively learns and leverages generalisable tabular representations, enabling efficient generation of high-quality synthetic tabular data, particularly with strong structural fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。