用元特征分析表格模型差异,发现现有方法难以稳定解释性能差距。
Explaining Tabular Foundation Model Differences Through Meta-Features
- 通过元特征关联模型性能差异,检验不同模型族在表格任务上的表现。
- 仅发现少数元特征与模型差距相关,且多数不具泛化能力。
- 适合关注模型选择与数据特性的研究人员参考。
随着表格基础模型的兴起,传统模型在许多任务中仍表现良好,如何为表格数据选择合适模型仍具挑战。本文基于TabArena基准测试结果,探究数据集元特征是否能解释模型族间的性能差距。经严格统计检验并控制错误发现率后发现:(1) 神经网络与树模型间的性能差异无法由任何元特征解释;(2) 非基础模型与基础模型之间的差距虽有一项关联稳健,但在留一数据集外预测中无法泛化;(3) TabICLv2与TabPFN-2.6间的差距有一项关联既稳健又提升留出集预测效果。此外,留一数据集外分析显示,元特征预测器未能显著优于简单基线。总体表明,表格数据高度异质,全局元特征方法不足以在51个TabArena数据集上提供可靠解释。
原文摘要 · Abstract (English)
With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular dataset remains difficult. We investigate whether dataset meta-features can explain performance gaps between model families on tabular prediction tasks. Using the TabArena benchmark results, we analyze dataset-level performance gaps and relate them to model-agnostic meta-features. After strict statistical tests with false discovery control, we find that (1) for neural network vs. tree gaps, no meta-feature survives false discovery control, (2) for non-foundation vs. foundation model gaps, one association is robust but does not generalize when tested in leave-one-dataset-out prediction, and (3) for TabICLv2 vs. TabPFN-2.6, one robust association also improves held-out prediction. Furthermore, we conduct a leave-one-dataset-out analysis and find that meta-feature predictors fail to improve meaningfully over a simple baseline. Overall, our results show the heterogeneity of tabular datasets and that global meta-feature approaches are not robust enough to offer explanations on the 51 TabArena datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。