针对不平衡数据,提出可证明有效的分类优化算法
Principled Algorithms for Optimizing Generalized Metrics in Binary Classification
- 将评估指标优化转化为代价敏感学习问题,设计新代理损失函数
- 算法在有限样本下有理论保证,优于传统阈值法
- 适合高不平衡或代价敏感场景的模型优化,如医疗诊断
在类别严重不平衡或代价不对称的应用中,Fβ度量、AM度量、杰卡德相似系数和加权准确率等指标比标准二分类损失更合适。然而,优化这些指标存在显著的计算与统计挑战。现有方法通常依赖贝叶斯最优分类器的刻画,采用先估计概率再找最优阈值的两阶段策略,导致算法不适用于受限假设集,且缺乏有限样本性能保证。本文提出原则性算法以优化广义指标,具备H-一致性与有限样本泛化界。方法将指标优化重构为广义代价敏感学习问题,设计出具有可证明H-一致性保证的新代理损失函数。基于此框架,开发了新算法METRO,具备强理论性能保障。实验结果表明,该方法在多个基准上优于已有基线。
原文摘要 · Abstract (English)
In applications with significant class imbalance or asymmetric costs, metrics such as the $F_β$-measure, AM measure, Jaccard similarity coefficient, and weighted accuracy offer more suitable evaluation criteria than standard binary classification loss. However, optimizing these metrics present significant computational and statistical challenges. Existing approaches often rely on the characterization of the Bayes-optimal classifier, and use threshold-based methods that first estimate class probabilities and then seek an optimal threshold. This leads to algorithms that are not tailored to restricted hypothesis sets and lack finite-sample performance guarantees. In this work, we introduce principled algorithms for optimizing generalized metrics, supported by $H$-consistency and finite-sample generalization bounds. Our approach reformulates metric optimization as a generalized cost-sensitive learning problem, enabling the design of novel surrogate loss functions with provable $H$-consistency guarantees. Leveraging this framework, we develop new algorithms, METRO (Metric Optimization), with strong theoretical performance guarantees. We report the results of experiments demonstrating the effectiveness of our methods compared to prior baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。