SOTA宣称常缺乏扎实证据,仅靠平均分领先未必代表真正优越。
Position: State-of-the-Art Claims Require State-of-the-Art Evidence
- 用平均得分排名判定SOTA存在漏洞,忽略实际性能一致性
- 超半数顶尖模型对比中,关键优势属性不成立,如效果量或鲁棒性
- 建议诚实报告结果,避免夸大宣称,提升模型比较可信度
当前人工智能与机器学习研究中,对最先进(SOTA)的宣称极为普遍。这类宣称依赖基准测试的聚合得分排名,但这种仅基于平均分的证据往往不足以支撑强结论。我们发现AI基准测试普遍存在‘宣称与证据不匹配’的问题。声称SOTA隐含了模型在多数任务上均显著优于其他方法,但平均分微小提升仅反映排名靠前,并不代表真实优势。分析来自公开排行榜的十个跨领域基准后发现,在超过一半的顶级模型比较中,至少一项被默认的优势属性不成立:包括显著的效果量、任务间的一致性,或对数据集移除的鲁棒性。这些看似显著的聚合增益,实际上常由少数异常数据集驱动。即使在任务众多的基准中,这种脆弱性依然存在。因此我们主张,宣称语言应与证据强度匹配,无需额外实验,只需如实呈现结果,从而实现更精准、可解释的模型对比。
原文摘要 · Abstract (English)
State-of-the-Art (SOTA) claims pervade Artificial Intelligence (AI) and Machine Learning (ML) research. These claims rest on benchmark evaluations, where models are ranked by aggregate scores across tasks. Public benchmarks or leaderboards are the most visible instance, but the same structure appears in paper tables throughout the literature. However, such minimal evidence often cannot support these strong claims. We identify a widespread claim-evidence gap in AI benchmarking. Claiming SOTA carries implicit assumptions beyond mean score superiority, suggesting that a model meaningfully outperforms alternatives across most tasks. However, a marginal improvement in the mean score merely indicates a top average rank rather than true superiority. Analyzing ten cross-domain benchmarks from public leaderboards, we found that in more than half of top-model comparisons, at least one commonly assumed property of superiority does not hold. These properties include meaningful effect size, consistency across tasks, or robustness to dataset removal. Instead, aggregate gains are frequently driven by outlier datasets. This fragility persists even in benchmarks with many tasks. We argue that claim language should reflect the strength of the underlying evidence. This requires no additional experiments, only honest reporting of what results actually show, enabling more precise and interpretable comparisons across models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。