通过混合合成先验提升表格模型泛化能力
Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
- 设计多样化合成数据先验,优化表格基础模型训练
- 在分类与回归任务中均超越现有最佳模型,且样本效率更高
- 适合关注表格数据预训练与少样本学习的研究者
自TabPFN问世以来,基于上下文学习(ICL)的表格基础模型(TFMs)挑战了机器学习领域长期存在的范式。这些模型仅在纯合成数据上预训练,无需接触真实数据,即可在多个数据集上实现出色泛化,通常仅需少量上下文示例。这促使研究重点从模型架构转向合成数据集的设计,即生成它们的先验分布。然而,先验设计的指导原则仍不明确。本文首次系统探究并识别出促进模型良好泛化的合成先验关键特性。基于此,我们提出Mitra,一种在精心筛选的、具备多样性、独特性且在真实表格数据上表现优异的合成先验混合数据上训练的表格基础模型。Mitra在分类与回归基准测试中持续优于当前最优的TFMs,如TabPFNv2和TabICL,且具备更优的样本效率。
原文摘要 · Abstract (English)
Since the seminal work of TabPFN, research on tabular foundation models (TFMs) based on in-context learning (ICL) has challenged long-standing paradigms in machine learning. Without seeing any real-world data, models pretrained on purely synthetic datasets generalize remarkably well across diverse datasets, often using only a moderate number of in-context examples. This shifts the focus in tabular machine learning from model architecture design to the design of synthetic datasets, or, more precisely, to the prior distributions that generate them. Yet the guiding principles for prior design remain poorly understood. This work marks the first attempt to address the gap. We systematically investigate and identify key properties of synthetic priors that allow pretrained TFMs to generalize well. Based on these insights, we introduce Mitra, a TFM trained on a curated mixture of synthetic priors selected for their diversity, distinctiveness, and performance on real-world tabular data. Mitra consistently outperforms state-of-the-art TFMs, such as TabPFNv2 and TabICL, across both classification and regression benchmarks, with better sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。