用特征值熵评估分类器,尤其适合不平衡数据。
The Eigenvalues Entropy as a Classifier Evaluation Measure
- 基于分类结果的特征值熵构建新评估指标。
- 在不平衡数据上表现优于传统指标如AUC、Gini。
- 可估计混淆矩阵,缓解类别不平衡问题。
分类是文本挖掘、手写字符识别、人脸识别、场景标注、计算机视觉和自然语言处理等众多实际应用中的关键技术。通常通过预测结果与训练集信息构建列联表,并借助评估度量来量化方法性能。现有许多评估指标在类别不平衡数据上表现不佳。本文提出使用特征值熵作为二分类或多分类问题的评估指标。针对二分类问题,给出了特征值与敏感性、特异性、ROC曲线下面积及吉尼指数之间的关系。本研究的一个副成果是能估计混淆矩阵,以应对类别不平衡带来的挑战。通过多个数据集验证,所提指标在性能上优于文献中的标准指标。
原文摘要 · Abstract (English)
Classification is a machine learning method used in many practical applications: text mining, handwritten character recognition, face recognition, pattern classification, scene labeling, computer vision, natural langage processing. A classifier prediction results and training set information are often used to get a contingency table which is used to quantify the method quality through an evaluation measure. Such measure, typically a numerical value, allows to choose a suitable method among several. Many evaluation measures available in the literature are less accurate for a dataset with imbalanced classes. In this paper, the eigenvalues entropy is used as an evaluation measure for a binary or a multi-class problem. For a binary problem, relations are given between the eigenvalues and some commonly used measures, the sensitivity, the specificity, the area under the operating receiver characteristic curve and the Gini index. A by-product result of this paper is an estimate of the confusion matrix to deal with the curse of the imbalanced classes. Various data examples are used to show the better performance of the proposed evaluation measure over the gold standard measures available in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。