用二维地图统一展示分类器所有评分排名,一目了然比高低。
The Tile: A 2D Map of Ranking Scores for Two-Class Classification
- 构建二维评分地图(Tile),整合准确率、召回率、F值等全部评价指标
- 可同时可视化多个分类器在无穷多评分下的排名关系
- 适合需综合评估分类器性能的研究者快速比较模型优劣
在计算机视觉与机器学习领域,新方法的严格评估至关重要。其中,分类器性能的比较与排序尤为关键,但现有工具如ROC和PR曲线仅基于两个分数,难以全面反映不同评分下的表现差异,也缺乏明确的排序能力。本文提出一种名为Tile的新工具,将两类分类器的所有排名分数(包括准确率、真正例率、正类预测值、Jaccard系数及所有F-beta分数)统一组织在一个二维地图中。我们研究了这些评分的性质,例如先验影响及与ROC空间的关系,并展示了如何通过与Tile对比来表征其他评分。结果表明,Tile能以单一可视化形式有效捕捉所有排名信息,实现性能解读与比较。
原文摘要 · Abstract (English)
In the computer vision and machine learning communities, as well as in many other research domains, rigorous evaluation of any new method, including classifiers, is essential. One key component of the evaluation process is the ability to compare and rank methods. However, ranking classifiers and accurately comparing their performances, especially when taking application-specific preferences into account, remains challenging. For instance, commonly used evaluation tools like Receiver Operating Characteristic (ROC) and Precision/Recall (PR) spaces display performances based on two scores. Hence, they are inherently limited in their ability to compare classifiers across a broader range of scores and lack the capability to establish a clear ranking among classifiers. In this paper, we present a novel versatile tool, named the Tile, that organizes an infinity of ranking scores in a single 2D map for two-class classifiers, including common evaluation scores such as the accuracy, the true positive rate, the positive predictive value, Jaccard's coefficient, and all F-beta scores. Furthermore, we study the properties of the underlying ranking scores, such as the influence of the priors or the correspondences with the ROC space, and depict how to characterize any other score by comparing them to the Tile. Overall, we demonstrate that the Tile is a powerful tool that effectively captures all the rankings in a single visualization and allows interpreting them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。