用真实数据微调表格大模型,显著提升预测准确率。
Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
- 用少量高质量真实数据继续预训练表格模型
- 在29个OpenML数据集上实现明显性能提升
- 适合需要高精度表格预测的应用场景
面向表格数据的基础模型(如TabPFN)在仅使用合成数据预训练时,在小数据集上表现优异。我们发现,通过有针对性的持续预训练阶段可显著提升其性能。具体而言,利用少量精心筛选的大规模真实世界数据集进行持续预训练,相比使用更广泛但可能含噪的数据源(如CommonCrawl或GitTables),能获得更好的下游预测准确性。我们提出的模型Real-TabPFN在OpenML自动机器学习基准的29个数据集上实现了显著性能提升。
原文摘要 · Abstract (English)
Foundation models for tabular data, like TabPFN, achieve strong performance on small datasets when pre-trained solely on synthetic data. We show that this performance can be significantly boosted by a targeted continued pre-training phase. Specifically, we demonstrate that leveraging a small, curated collection of large, real-world datasets for continued pre-training yields superior downstream predictive accuracy compared to using broader, potentially noisier corpora like CommonCrawl or GitTables. Our resulting model, Real-TabPFN, achieves substantial performance gains on 29 datasets from the OpenML AutoML Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。