针对海量少数类数据的极端不平衡分类问题,提出新损失函数提升模型鲁棒性。
Ultra-imbalanced classification guided by statistical information
- 从统计信息角度重新定义极端不平衡学习,提出全新建模范式
- 新损失函数在无限样本下仍能有效抵抗数据失衡,实验验证高效
- 特别适合金融反欺诈等工业场景中少数类样本极多的现实问题
真实世界分类任务中常遇到数据不平衡问题。以往研究主要关注少数类样本极少的情况,但实际中少数类样本丰富的极端不平衡现象更常见,如金融风控中的欺诈检测。本文提出一种全新的建模范式——超不平衡分类(Ultra-imbalanced Classification, UIC),突破传统以样本量为基准的不平衡认知。在UIC框架下,即使拥有无穷多训练样本,损失函数的表现仍存在本质差异。基于信息论思想,构建了通过统计信息比较不同损失函数的分析框架,并设计出可调增强损失(Tunable Boosting Loss)。该损失函数在理论上对数据不平衡具有鲁棒性,且在公开和工业数据集上的大量实验中表现出优异的实证效率。
原文摘要 · Abstract (English)
Imbalanced data are frequently encountered in real-world classification tasks. Previous works on imbalanced learning mostly focused on learning with a minority class of few samples. However, the notion of imbalance also applies to cases where the minority class contains abundant samples, which is usually the case for industrial applications like fraud detection in the area of financial risk management. In this paper, we take a population-level approach to imbalanced learning by proposing a new formulation called \emph{ultra-imbalanced classification} (UIC). Under UIC, loss functions behave differently even if infinite amount of training samples are available. To understand the intrinsic difficulty of UIC problems, we borrow ideas from information theory and establish a framework to compare different loss functions through the lens of statistical information. A novel learning objective termed Tunable Boosting Loss is developed which is provably resistant against data imbalance under UIC, as well as being empirically efficient verified by extensive experimental studies on both public and industrial datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。