arXiv:2412.04166cs.LGcs.NA2024-12

评估多分类模型的误判风险,提出无需依赖数据分布的新方法。

An In-Depth Examination of Risk Assessment in Multi-Class Classification Algorithms

  • 用置信区间生成技术改进风险评估,不依赖具体模型和数据分布。
  • 在多种模型和数据集上验证,新方法表现稳定且可靠。
  • 适合医疗、工程等高风险场景中需准确判断误判概率的用户。

高级分类算法正越来越多地应用于医疗、工程等安全关键领域。在这些应用中,机器学习模型的误分类可能导致重大财务或健康损失。为更好预测和应对此类损失,算法使用者需要估计模型对样本误分类的概率,这一任务称为风险评估。本文针对多种模型和数据集,数值分析了不同方法在解决风险评估问题上的表现。考虑两种解决方案:a) 校准技术,通过校准分类模型的输出概率以获得准确的概率输出;b) 基于合取预测(conformal prediction)的新型方法。该方法不依赖模型类型和数据分布,实现简单,并在多种应用场景中表现出合理性能。我们在广泛的模型和数据集上对比了不同方法的表现。

原文摘要 · Abstract (English)

Advanced classification algorithms are being increasingly used in safety-critical applications like health-care, engineering, etc. In such applications, miss-classifications made by ML algorithms can result in substantial financial or health-related losses. To better anticipate and prepare for such losses, the algorithm user seeks an estimate for the probability that the algorithm miss-classifies a sample. We refer to this task as the risk-assessment. For a variety of models and datasets, we numerically analyze the performance of different methods in solving the risk-assessment problem. We consider two solution strategies: a) calibration techniques that calibrate the output probabilities of classification models to provide accurate probability outputs; and b) a novel approach based upon the prediction interval generation technique of conformal prediction. Our conformal prediction based approach is model and data-distribution agnostic, simple to implement, and provides reasonable results for a variety of use-cases. We compare the different methods on a broad variety of models and datasets.

风险评估合取预测多分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。