arXiv:2409.05598stat.MLcond-mat.dis-nn2024-09被引 2

研究类别不平衡下重采样是否提升特征学习,发现有时不重采样反而更好。

When resampling/reweighting improves feature learning in imbalanced classification?: A toy-model study

  • 用简化模型分析类别不平衡时重采样对特征学习的影响。
  • 发现特定条件下不重采样反而性能最优,与已有实验一致。
  • 揭示损失函数对称性是关键,适合关注理论机制的研究者阅读。

本文通过一个二分类的简化模型,研究在类别不平衡情况下,类别级重采样/重加权对特征学习性能的影响。分析中采用高维输入空间极限,并保持数据集大小与输入维度之比为有限值,使用统计力学中的非严格复制方法。结果表明,在某些情形下,无论损失函数或分类器如何选择,不进行重采样/重加权反而能获得最佳特征学习性能,这一结论与Cao等(2019)和Kang等(2019)的近期发现一致。研究进一步揭示,该现象的关键在于损失函数与问题设置的对称性。基于此,我们提出一个在多分类设定下具有相同性质的更简化模型。这些工作澄清了类别级重采样/重加权在类别不平衡分类中何时真正有效。

原文摘要 · Abstract (English)

A toy model of binary classification is studied with the aim of clarifying the class-wise resampling/reweighting effect on the feature learning performance under the presence of class imbalance. In the analysis, a high-dimensional limit of the input space is taken while keeping the ratio of the dataset size against the input dimension finite and the non-rigorous replica method from statistical mechanics is employed. The result shows that there exists a case in which the no resampling/reweighting situation gives the best feature learning performance irrespectively of the choice of losses or classifiers, supporting recent findings in Cao et al. (2019); Kang et al. (2019). It is also revealed that the key of the result is the symmetry of the loss and the problem setting. Inspired by this, we propose a further simplified model exhibiting the same property in the multiclass setting. These clarify when the class-wise resampling/reweighting becomes effective in imbalanced classification.

类别不平衡特征学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。