为性能排名建立通用理论,统一评估标准与应用偏好。
Foundations of the Theory of Performance-Based Ranking
- 基于概率与序理论构建性能数学框架
- 提出首个满足公理的性能排序方法,涵盖常见评估指标
- 可灵活适配不同应用场景的优先级需求
根据性能对算法、设备、方法或模型进行排序,同时考虑特定应用的偏好,是一项挑战。为此,本文建立了性能排名的通用理论基础。首先,提出一个严谨的框架,基于概率与序理论,能够(1)将性能视为数学对象处理,(2)表达性能之间的优劣或等价关系,(3)通过满意度变量建模任务,(4)考虑评估属性,(5)定义得分,(6)通过重要性变量表达应用特定偏好。在此基础上,首次提出性能排序的公理化定义。接着,引入一个通用参数化得分族——排名得分,可用于构建满足公理的排名,并融合应用偏好。最后,在二分类场景下,证明该得分族包含准确率、真正率(召回率、敏感度)、真负率(特异度)、正预测值(精确度)和F1等经典指标;但同时也表明,某些常用分类器比较指标不满足该公理体系。
原文摘要 · Abstract (English)
Ranking entities such as algorithms, devices, methods, or models based on their performances, while accounting for application-specific preferences, is a challenge. To address this challenge, we establish the foundations of a universal theory for performance-based ranking. First, we introduce a rigorous framework built on top of both the probability and order theories. Our new framework encompasses the elements necessary to (1) manipulate performances as mathematical objects, (2) express which performances are worse than or equivalent to others, (3) model tasks through a variable called satisfaction, (4) consider properties of the evaluation, (5) define scores, and (6) specify application-specific preferences through a variable called importance. On top of this framework, we propose the first axiomatic definition of performance orderings and performance-based rankings. Then, we introduce a universal parametric family of scores, called ranking scores, that can be used to establish rankings satisfying our axioms, while considering application-specific preferences. Finally, we show, in the case of two-class classification, that the family of ranking scores encompasses well-known performance scores, including the accuracy, the true positive rate (recall, sensitivity), the true negative rate (specificity), the positive predictive value (precision), and F1. However, we also show that some other scores commonly used to compare classifiers are unsuitable to derive performance orderings satisfying the axioms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。