arXiv:2603.20587cs.LGcs.IT2026-03被引 1

揭示高类别数下神经网络特征的几何结构规律。

Neural collapse in the orthoplex regime

  • 在维度满足d+2≤n≤2d时,分析特征向量的几何形态。
  • 发现特征收敛至正交单纯形(orthoplex)结构。
  • 适用于理解大类别分类模型的内在机制。

当训练神经网络进行分类时,若特征空间维数d与类别数n满足n≤d+1,训练集的特征向量会坍缩到正则单纯形的顶点,这一现象称为神经坍缩。对于语言模型等n≫d的应用场景,神经坍缩仍会发生,但呈现不同的几何结构。本文研究d+2≤n≤2d时的“正交单纯形区域”,利用Radon定理和凸性分析,刻画了该条件下特征向量的几何形态。结果表明,特征向量趋于形成正交单纯形结构,揭示了高类别数下的统一几何规律。

原文摘要 · Abstract (English)

When training a neural network for classification, the feature vectors of the training set are known to collapse to the vertices of a regular simplex, provided the dimension $d$ of the feature space and the number $n$ of classes satisfies $n\leq d+1$. This phenomenon is known as neural collapse. For other applications like language models, one instead takes $n\gg d$. Here, the neural collapse phenomenon still occurs, but with different emergent geometric figures. We characterize these geometric figures in the orthoplex regime where $d+2\leq n\leq 2d$. The techniques in our analysis primarily involve Radon's theorem and convexity.

神经坍缩几何结构分类模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。