提出可证明有效的不平衡数据学习框架与新算法,提升分类准确率。
Balancing the Scales: A Theoretical and Algorithmic Framework for Learning from Imbalanced Data
- 构建基于边际的理论框架,解决不平衡分类的泛化问题
- 新损失函数实现强一致性,理论保证优于传统方法
- 算法IMMAX适用于多种模型,适合长尾分布场景
类别不平衡仍是机器学习中的重大挑战,尤其在具有长尾分布的多分类任务中。现有方法如数据重采样、代价敏感和逻辑损失调整虽常用且有效,但缺乏坚实的理论基础。本文证明代价敏感方法不满足贝叶斯一致性。为此,提出一种新的理论框架,用于分析不平衡分类中的泛化性能。设计适用于二分类和多分类的类别不平衡边际损失函数,证明其强H-一致性,并基于经验损失和一种新的类敏感Rademacher复杂度,推导出学习保证。基于这些理论结果,提出通用的新算法IMMAX(不平衡边际最大化),引入置信度边界,适用于多种假设集。尽管聚焦理论,仍通过大量实验验证算法优于现有基线。
原文摘要 · Abstract (English)
Class imbalance remains a major challenge in machine learning, especially in multi-class problems with long-tailed distributions. Existing methods, such as data resampling, cost-sensitive techniques, and logistic loss modifications, though popular and often effective, lack solid theoretical foundations. As an example, we demonstrate that cost-sensitive methods are not Bayes-consistent. This paper introduces a novel theoretical framework for analyzing generalization in imbalanced classification. We propose a new class-imbalanced margin loss function for both binary and multi-class settings, prove its strong $H$-consistency, and derive corresponding learning guarantees based on empirical loss and a new notion of class-sensitive Rademacher complexity. Leveraging these theoretical results, we devise novel and general learning algorithms, IMMAX (Imbalanced Margin Maximization), which incorporate confidence margins and are applicable to various hypothesis sets. While our focus is theoretical, we also present extensive empirical results demonstrating the effectiveness of our algorithms compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。