提出评估分类器一致性值显著性的通用方法,解决现有标准主观随意的问题。
Significativity Indices for Agreement Values
- 基于有限数据集和概率分布,构建两种一致性显著性指数
- 提供高效算法,克服计算复杂度难题
- 适用于医学与人工智能中的分类器性能评估
一致性度量(如Cohen's kappa、组内相关系数)用于衡量两个或多个分类器之间的匹配程度,广泛应用于医学治疗效果评估及人工智能中分类器简化后的近似度量化。通过比较各分类器与黄金标准的一致性值大小,可直观判断其优劣。然而,仅凭一致性的数值判断方法好坏,需依赖标准化的显著性指标。现有文献中针对Cohen's kappa的质量分级大多过于简单,边界设定主观。本文提出一种通用框架,用于评估任意一致性值的显著性,并引入两种显著性指数:一种适用于有限样本数据集,另一种处理分类概率分布情形。同时,针对该指标的计算挑战,提出若干高效算法以实现快速评估。
原文摘要 · Abstract (English)
Agreement measures, such as Cohen's kappa or intraclass correlation, gauge the matching between two or more classifiers. They are used in a wide range of contexts from medicine, where they evaluate the effectiveness of medical treatments and clinical trials, to artificial intelligence, where they can quantify the approximation due to the reduction of a classifier. The consistency of different classifiers to a golden standard can be compared simply by using the order induced by their agreement measure with respect to the golden standard itself. Nevertheless, labelling an approach as good or bad exclusively by using the value of an agreement measure requires a scale or a significativity index. Some quality scales have been proposed in the literature for Cohen's kappa, but they are mainly naïve, and their boundaries are arbitrary. This work proposes a general approach to evaluate the significativity of any agreement value between two classifiers and introduces two significativity indices: one dealing with finite data sets, the other one handling classification probability distributions. Moreover, this manuscript addresses the computational challenges of evaluating such indices and proposes some efficient algorithms for their evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。