深度网络先学大类,后学细类,揭示分类学习的层次演化机制。
Hypernym Bias: Unraveling Deep Classifier Training Dynamics through the Lens of Class Hierarchy
- 通过追踪特征流形变化,发现网络按层级逐步学习类别关系。
- 高层语义(超词)类别的神经坍缩现象早于细类出现。
- 适合研究模型学习动态与语义结构对齐的研究者阅读。
我们通过分析类别间层级关系在训练过程中的演变,研究深度分类器的训练动态。大量实验表明,分类学习可理解为标签聚类过程:网络在训练初期优先区分高层次(超词)类别,后期才细化具体(下位词)类别。本文提出新框架,追踪训练中特征流形的演化,揭示类别层级结构如何在各层网络中逐步形成并优化。分析显示,学习到的表示与数据集语义结构高度一致,量化描述了聚类过程。值得注意的是,在超词标签空间中,神经坍缩特性出现时间早于下位词空间,有助于弥合学习初期与末期的差距。研究为理解深度网络中的层次学习机制提供了新视角,推动对深度学习动态的进一步认识。
原文摘要 · Abstract (English)
We investigate the training dynamics of deep classifiers by examining how hierarchical relationships between classes evolve during training. Through extensive experiments, we argue that the learning process in classification problems can be understood through the lens of label clustering. Specifically, we observe that networks tend to distinguish higher-level (hypernym) categories in the early stages of training, and learn more specific (hyponym) categories later. We introduce a novel framework to track the evolution of the feature manifold during training, revealing how the hierarchy of class relations emerges and refines across the network layers. Our analysis demonstrates that the learned representations closely align with the semantic structure of the dataset, providing a quantitative description of the clustering process. Notably, we show that in the hypernym label space, certain properties of neural collapse appear earlier than in the hyponym label space, helping to bridge the gap between the initial and terminal phases of learning. We believe our findings offer new insights into the mechanisms driving hierarchical learning in deep networks, paving the way for future advancements in understanding deep learning dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。