arXiv:2504.09759cs.LG2025-04

用游戏评分系统重新评估分类器,更公平地衡量真实能力与鲁棒性。

Enhancing Classifier Evaluation: A Fairer Benchmarking Strategy Based on Ability and Robustness

  • 结合项目反应理论与游戏评级系统,评估模型在难题上的表现能力。
  • 仅15%数据集真正困难,减少50%数据集仍可提供相近评估效果。
  • 随机森林能力得分最高,适合关注模型真实泛化能力的研究者。

机器学习中的基准测试常忽略数据集复杂度与算法泛化能力的双重考量,导致评估偏向简单样本上表现好的模型。本文提出一种新方法,融合项目反应理论(IRT)与原用于游戏玩家评级的Glicko-2系统,通过模拟分类器之间的对抗赛,动态更新其能力评分(含等级、离散度与波动性)。该方法能更公正地反映模型在难题上的真实表现。以OpenML-CC18为例,研究发现仅15%的数据集具有真正挑战性,而缩减至原数据集50%时仍具备相当的评估效能。实验中,随机森林在所有算法中获得最高能力评分。结果表明,优化基准设计应聚焦数据质量,并采用兼顾难度与模型水平的评估策略。

原文摘要 · Abstract (English)

Benchmarking is a fundamental practice in machine learning (ML) for comparing the performance of classification algorithms. However, traditional evaluation methods often overlook a critical aspect: the joint consideration of dataset complexity and an algorithm's ability to generalize. Without this dual perspective, assessments may favor models that perform well on easy instances while failing to capture their true robustness. To address this limitation, this study introduces a novel evaluation methodology that combines Item Response Theory (IRT) with the Glicko-2 rating system, originally developed to measure player strength in competitive games. IRT assesses classifier ability based on performance over difficult instances, while Glicko-2 updates performance metrics - such as rating, deviation, and volatility - via simulated tournaments between classifiers. This combined approach provides a fairer and more nuanced measure of algorithm capability. A case study using the OpenML-CC18 benchmark showed that only 15% of the datasets are truly challenging and that a reduced subset with 50% of the original datasets offers comparable evaluation power. Among the algorithms tested, Random Forest achieved the highest ability score. The results highlight the importance of improving benchmark design by focusing on dataset quality and adopting evaluation strategies that reflect both difficulty and classifier proficiency.

分类器评估基准测试模型能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。