arXiv:2604.06814cs.LGcs.AI2026-04被引 2

大规模对比树模型、神经网络和大模型在表格数据上的表现,发现无绝对优胜者。

OmniTabBench: Mapping the Empirical Frontiers of GBDTs, Neural Networks, and Foundation Models for Tabular Data at Scale

  • 构建3030个数据集的超大规模基准测试集
  • 揭示不同数据特征下各类模型的适用条件
  • 用大模型分类数据领域,减少评估偏差

尽管传统的树模型长期主导表格数据任务,深度神经网络和新兴基础模型已对其构成挑战,但尚无统一公认的最优范式。现有基准通常仅含少于100个数据集,存在评估不足与选择偏差的担忧。为此,我们提出OmniTabBench,目前最大的表格数据基准,涵盖3030个数据集,覆盖多样任务,通过大语言模型从多元来源收集并按行业分类。我们在该基准上对各模型家族的前沿模型进行前所未有的大规模实证评估,确认不存在占优模型。进一步通过解耦的元特征分析(如数据集规模、特征类型、特征与目标偏度/峰度),揭示特定模型类别在何种条件下更优,为实际应用提供比以往综合指标研究更清晰、更具操作性的指导。

原文摘要 · Abstract (English)

While traditional tree-based ensemble methods have long dominated tabular tasks, deep neural networks and emerging foundation models have challenged this primacy, yet no consensus exists on a universally superior paradigm. Existing benchmarks typically contain fewer than 100 datasets, raising concerns about evaluation sufficiency and potential selection biases. To address these limitations, we introduce OmniTabBench, the largest tabular benchmark to date, comprising 3030 datasets spanning diverse tasks that are comprehensively collected from diverse sources and categorized by industry using large language models. We conduct an unprecedented large-scale empirical evaluation of state-of-the-art models from all model families on OmniTabBench, confirming the absence of a dominant winner. Furthermore, through a decoupled metafeature analysis, which examines individual properties such as dataset size, feature types, feature and target skewness/kurtosis, we elucidate conditions favoring specific model categories, providing clearer, more actionable guidance than prior compound-metric studies.

表格数据大模型基准测试模型比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。