arXiv:2601.12011cs.LG2026-01被引 1

早停时加权损失能平衡类别学习,避免少数类被忽视。

Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features

  • 用简化模型揭示损失加权如何改变学习顺序。
  • 早期训练中加权使多数与少数类别同时学习。
  • 适合关注类别不平衡问题的从业者参考。

在现代深度学习中,损失重加权的应用呈现出复杂图景。尽管在过参数化神经网络于高维数据上训练时,它无法改变最终的学习阶段,但实证证据一致表明其在训练早期具有显著优势。为清晰展示并分析此现象,我们引入一个小规模模型(SSM)。该模型专门设计用于抽象深度神经网络架构和输入数据的内在复杂性,同时保留谱成分中不平衡结构的关键信息。一方面,SSM显示,标准经验风险最小化在训练初期优先学习区分多数类,从而延迟少数类的学习;另一方面,重加权恢复了平衡的学习动态,使与多数类和少数类相关的特征得以同时学习。

原文摘要 · Abstract (English)

The application of loss reweighting in modern deep learning presents a nuanced picture. While it fails to alter the terminal learning phase in overparameterized deep neural networks (DNNs) trained on high-dimensional datasets, empirical evidence consistently shows it offers significant benefits early in training. To transparently demonstrate and analyze this phenomenon, we introduce a small-scale model (SSM). This model is specifically designed to abstract the inherent complexities of both the DNN architecture and the input data, while maintaining key information about the structure of imbalance within its spectral components. On the one hand, the SSM reveals how vanilla empirical risk minimization preferentially learns to distinguish majority classes over minorities early in training, consequently delaying minority learning. In stark contrast, reweighting restores balanced learning dynamics, enabling the simultaneous learning of features associated with both majorities and minorities.

损失加权类别不平衡训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。