发现模型泛化依赖平坦性,而非神经坍缩。
Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
- 用'领悟'训练模式分离泛化与训练过程,追踪动态变化。
- 只有平坦性在泛化前稳定出现,神经坍缩无关紧要。
- 适合研究泛化机制的学者,尤其关注损失曲面几何者。
神经坍缩(即类间表示高度对称聚集)常被视作泛化的标志,而损失曲面的平坦性也被理论和实证关联到泛化能力。但二者是否为泛化前提仍不明。本文通过'领悟'(grokking)这一训练范式——记忆先于泛化——实现泛化与训练动态的时间分离。结果发现,尽管神经坍缩与相对平坦性均在泛化初期出现,仅平坦性能一致预测泛化发生。强制或阻止坍缩的模型泛化表现相当;而远离平坦解的模型则延迟泛化,甚至在非典型场景中表现出类似'领悟'的行为。理论上,在经典假设下,神经坍缩可导出相对平坦性,解释了二者的共现。结论支持:相对平坦性是泛化更基础且可能必要的性质,而'领悟'可作为探测其几何本质的强大工具。
原文摘要 · Abstract (English)
Neural collapse, i.e., the emergence of highly symmetric, class-wise clustered representations, is frequently observed in deep networks and is often assumed to reflect or enable generalization. In parallel, flatness of the loss landscape has been theoretically and empirically linked to generalization. Yet, the causal role of either phenomenon remains unclear: Are they prerequisites for generalization, or merely by-products of training dynamics? We disentangle these questions using grokking, a training regime in which memorization precedes generalization, allowing us to temporally separate generalization from training dynamics and we find that while both neural collapse and relative flatness emerge near the onset of generalization, only flatness consistently predicts it. Models encouraged to collapse or prevented from collapsing generalize equally well, whereas models regularized away from flat solutions exhibit delayed generalization, resembling grokking, even in architectures and datasets where it does not typically occur. Furthermore, we show theoretically that neural collapse leads to relative flatness under classical assumptions, explaining their empirical co-occurrence. Our results support the view that relative flatness is a potentially necessary and more fundamental property for generalization, and demonstrate how grokking can serve as a powerful probe for isolating its geometric underpinnings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。