研究类别不平衡下重采样是否提升特征学习,发现有时不重采样反而更好。
When resampling/reweighting improves feature learning in imbalanced classification?: A toy-model study
- 用简化模型分析类别不平衡时重采样对特征学习的影响。
- 发现特定条件下不重采样反而性能最优,与已有实验一致。
- 揭示损失函数对称性是关键,适合关注理论机制的研究者阅读。
本文通过一个二分类的简化模型,研究在类别不平衡情况下,类别级重采样/重加权对特征学习性能的影响。分析中采用高维输入空间极限,并保持数据集大小与输入维度之比为有限值,使用统计力学中的非严格复制方法。结果表明,在某些情形下,无论损失函数或分类器如何选择,不进行重采样/重加权反而能获得最佳特征学习性能,这一结论与Cao等(2019)和Kang等(2019)的近期发现一致。研究进一步揭示,该现象的关键在于损失函数与问题设置的对称性。基于此,我们提出一个在多分类设定下具有相同性质的更简化模型。这些工作澄清了类别级重采样/重加权在类别不平衡分类中何时真正有效。
原文摘要 · Abstract (English)
A toy model of binary classification is studied with the aim of clarifying the class-wise resampling/reweighting effect on the feature learning performance under the presence of class imbalance. In the analysis, a high-dimensional limit of the input space is taken while keeping the ratio of the dataset size against the input dimension finite and the non-rigorous replica method from statistical mechanics is employed. The result shows that there exists a case in which the no resampling/reweighting situation gives the best feature learning performance irrespectively of the choice of losses or classifiers, supporting recent findings in Cao et al. (2019); Kang et al. (2019). It is also revealed that the key of the result is the symmetry of the loss and the problem setting. Inspired by this, we propose a further simplified model exhibiting the same property in the multiclass setting. These clarify when the class-wise resampling/reweighting becomes effective in imbalanced classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。