提出最小模型解析表征坍缩机制,揭示错误样本如何引发崩溃。
A Minimal Model of Representation Collapse: Frustration, Stop-Gradient, and Dynamics
- 用嵌入模型分析梯度流动态,闭式求解固定点
- 少量无法分类的样本诱发坍缩,出现缓慢时间尺度
- 停止梯度可稳定分类间隔,适合研究自监督学习
自监督表示学习能从无标签数据中提取结构化隐特征,实现跨任务和领域鲁棒迁移。但常遭遇表征坍缩——嵌入失去区分性,不同输入变得不可分辨。为理解其机制与防坍缩要素,我们构建一个仅含嵌入的最小模型,通过分类-表示设置量化坍缩,以标签-嵌入几何收缩衡量。当数据完全可分类时模型不坍缩;少量无法一致分类的样本会通过早期性能提升后的慢时间尺度引发坍缩。在同一框架下,引入共享投影头并应用训练动态层面的停止梯度,分析其固定点并建立动力学平均场自洽描述,表明停止梯度可实现非坍缩解,在有扰动时仍保持有限分类分离。进一步在线性教师-学生模型中验证了相同定性动态与防坍缩效应,说明该最小理论捕捉了超越纯嵌入设定的关键特征。
原文摘要 · Abstract (English)
Self-supervised representation learning is central to modern machine learning because it extracts structured latent features from unlabeled data and enables robust transfer across tasks and domains. However, it can suffer from representation collapse, a widely observed failure mode in which embeddings lose discriminative structure and distinct inputs become indistinguishable. To understand the mechanisms that drive collapse and the ingredients that prevent it, we introduce a minimal embedding-only model whose gradient-flow dynamics and fixed points can be analyzed in closed form, using a classification-representation setting as a concrete playground where collapse is directly quantified through the contraction of label-embedding geometry. We illustrate that the model does not collapse when the data are perfectly classifiable, while a small fraction of frustrated samples that cannot be classified consistently induces collapse through an additional slow time scale that follows the early performance gain. Within the same framework, we examine collapse prevention by adding a shared projection head and applying stop-gradient at the level of the training dynamics. We analyze the resulting fixed points and develop a dynamical mean-field style self-consistency description, showing that stop-gradient enables non-collapsed solutions and stabilizes finite class separation under frustration. We further verify empirically that the same qualitative dynamics and collapse-prevention effects appear in a linear teacher-student model, indicating that the minimal theory captures features that persist beyond the pure embedding setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。