现有基准评估忽略模型在特定数据集上的不可替代性,导致真正强的模型被低估。
Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths

- 提出以数据集级最优表现前沿为新评估视角,判断模型是否不可替代
- 发现主流平均指标与模型不可替代性相关性低,掩盖了独特优势
- 建议用能否提升各数据集峰值性能来衡量模型进步,而不仅是平均表现
表格机器学习基准通常通过跨数据集的分数、排名或两两胜率平均值来总结性能。这类聚合指标对选择稳健的默认模型有用,但可能掩盖另一个问题:哪些模型对实现特定数据集的最高性能是必需的?我们主张基准评估也应关注数据中心的峰值性能前沿,即每个数据集上统计上可支持的最佳表现。在此框架下,模型可能表现为不可替代、充分、冗余或不可靠,取决于其相对于其他模型在前沿中的位置。在TabArena基准上应用该框架发现,常见聚合指标高度相关,主要衡量一致性与避免失败,但与数据集级不可替代性关联较弱。因此,那些在多个数据集中表现尚可但从不夺冠的模型被奖励,而具有独特数据集优势的模型在聚合评价中却显得平庸。故此,基准进展不仅应看聚合指标提升,更要看新模型是否拓展了各数据集可达成的峰值性能集合。
原文摘要 · Abstract (English)
Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but they can obscure a different question: which models are necessary to attain peak performance on particular datasets? We argue that benchmark evaluation should also consider the data-centric peak performance frontier, defined by the best statistically supported performance achieved on each dataset. From this perspective, a model may be irreplaceable, sufficient, redundant, or fallible depending on where it lies on the frontier relative to other models. Applying this framework to the TabArena benchmark, we find that common aggregation metrics are highly correlated and largely measure consistency and avoiding failures, while being much less aligned with dataset-level irreplaceability. Consequently, models performing decently across datasets without ever being the best choice are rewarded while models with unique dataset-specific strengths appear mediocre under aggregation. Hence, benchmark progress should be measured not only by improvements on aggregation metrics but also by whether new models expand the set of attainable peak performances across datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。