统一分类性能指标尺度,让不同数据集上的模型表现可直接比较。
Introducing the O-Value: A Universal Standardization for Confusion-Matrix-Based Classification Performance Metrics
- 提出o值标准化方法,将各类分类指标映射到0~1统一量纲。
- o值代表实际性能在随机表现分布中的百分位排名,解释清晰。
- 适用于各种场景,尤其适合不平衡数据下的模型对比与监控。
现有分类性能指标因量纲不同且对类别不平衡敏感,难以跨数据集比较。本文提出一种通用标准化方法——出性能标准化(OPS),将任意基于混淆矩阵的分类性能指标映射至[0,1]区间。所得的o值表示实际性能在参考表现分布中的百分位排名,具有统一且直观的解释。该框架使不同不平衡率测试集间的性能评估、比较与监控成为可能。我们在多个真实世界数据集上验证了o值在多种常用指标上的适用性与稳健性,展示了其在多类分类任务中的实用价值。
原文摘要 · Abstract (English)
Many classification performance metrics exist, each suited to a specific application. However, these metrics often differ in scale and can exhibit varying sensitivity to class imbalance rates in the test set. As a result, it is difficult to use the nominal values of these metrics to evaluate, compare and monitor classification performances, especially when imbalance rates vary. To address this problem, we introduce the outperformance standardization (OPS) function, a universal standardization method for confusion-matrix-based classification performance (CMBCP) metrics. It maps any given metric to a common scale of $[0,1]$, while providing a clear and consistent interpretation. Specifically, the resulting OPS value (o-value) represents the percentile rank of the observed classification performance within a reference distribution of possible performances. This unified framework enables meaningful comparison and monitoring of classification performance across test sets with differing imbalance rates. We illustrate how o-values can be applied to a variety of commonly used classification performance metrics and demonstrate the utility and robustness of our method through experiments on real-world datasets spanning multiple classification applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。