改进模型评估方法,让排名更真实反映性能差距。
MARS: Magnitude-Aware Rank Statistics
- 引入相对差距系数加权排名,解决传统方法忽略性能差距的问题。
- 通过动态投影处理极端情况,使统计结果更稳定可靠。
- 适合需要精细比较多个模型性能的研究者使用。
机器学习模型的全面评估是确保其稳健性和一致性的关键。为总结实验结果并选出优胜模型,通常采用临界差异(CD)图。然而标准CD图依赖离散排名,忽略了模型间性能差距的大小,存在“忽视幅度”的问题。为此,我们提出幅度感知排名统计(MARS),将相对边际系数作为权重作用于离散排名,该系数根据最优与最差模型间的距离动态缩放排名,并引入动态投影机制处理边界情况。在计算CD值后,MARS能更真实地呈现模型性能差异,在大规模实验设置中提供更深入的洞察。
原文摘要 · Abstract (English)
Comprehensive evaluation of machine learning models is the key to make sure that they perform as robustly and consistently as desired. In order to summarize the experimental results and pick a winner, Critical Difference (CD) diagrams are used. Standard CD diagrams rely on discrete ranks, discarding the magnitude of performance gaps between models, raising an issue which we call magnitude-blindness. In order to address this issue, we propose Magnitude-Aware Rank Statistics (MARS) that incorporates a relative margin coefficient as a weight for the discrete ranks. This coefficient scales ranks based on the distance between the best and worst performers, with a dynamic projection to handle boundary cases. Followed by the calculation of a CD value, MARS results in a more realistic statistical representation of differences of model performances and more insights on how methods actually perform in vast and extensive experimental settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。