arXiv:2606.30410cs.LGcs.AI2026-06被引 4

首个统一基准测试揭示表格模型在非独立同分布数据上表现有限。

Beyond IID: How General Are Tabular Foundation Models, Really?

论文配图:Beyond IID: How General Are Tabular Foundation Models, Really?
图 1 · 摘自论文原文
  • 构建统一基准BeyondArena,覆盖多种任务与数据规模
  • 11个模型在非独立同分布数据上表现远逊于传统模型
  • 适合关注真实场景泛化能力的研究者

面向表格数据的预测机器学习基础模型近年在学术界和工业界广受关注。然而,当前跨领域评估仍分散且不透明,研究者多依赖标准基准,而这些基准主要针对模型已擅长的独立同分布(IID)数据,忽视了更具挑战性的场景。为此,我们提出BeyondArena——首个统一的表格数据综合基准,支持独立同分布、时间序列、分组等多样任务,覆盖不同样本量与特征维度,包含文本与高基数特征,来自多学科数据集。为实现统一评估,我们还推出了Data Foundry框架与元数据规范,用于表格数据集的标准化整理。在142个精选数据集上对11种模型的实验表明,现有基础模型仅在小到中等规模的IID数据上表现优异,而在非独立同分布、大规模与高维数据上,传统树模型与深度学习模型仍占主导。BeyondArena为表格基础模型研究指明了更具挑战性的方向。

原文摘要 · Abstract (English)

Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Research communities across disciplines are increasingly evaluating tabular foundation models on diverse datasets and tasks. However, these task- and discipline-specific evaluations remain largely inaccessible to model researchers because benchmark software and evaluation protocols are fragmented. As a result, model researchers rely on standard benchmarks, which are mostly defined for tasks where tabular foundation models already excel. The most challenging scenarios are excluded, limiting meaningful progress in the field by focusing on marginal improvements on IID data rather than on broader, more demanding challenges. To overcome this, we introduce BeyondArena, the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types (with text, with high cardinality) from a broad range of disciplines. To enable unified benchmarking beyond standard benchmarks, we introduce Data Foundry, a Python framework and metadata schema for curating tabular datasets for predictive machine learning. Our results across 11 models and 142 curated datasets show that existing tabular foundation models excel on tiny- to medium-sized IID data, while traditional tree-based and deep learning models still dominate on non-IID, large, and high-dimensional datasets. BeyondArena guides model research for the most demanding challenges in tabular data, enabling progress towards truly foundational tabular models.

表格数据基础模型泛化能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。