arXiv:2608.15632cs.LGcond-mat.dis-nn2026-08

发现跨模态分类的神经表示具有稀疏中心对齐结构,可预测模型性能。

Sparse Prototype Code Underlies Classification and Prediction Across Modalities

论文配图:Sparse Prototype Code Underlies Classification and Prediction Across Modalities
图 1 · 摘自论文原文
  • 基于类别中心与竞争类中心的几何相关性,构建解析理论。
  • 理论准确预测多模态模型分类精度,且随模型规模提升而改善。
  • 仅需少数关键中心坐标即可高精度预测,契合稀疏特征提取思路。

神经表示已成为研究现代AI模型内部机制的核心工具,但其复杂的高维结构难以解释。我们发现,视觉、音频和语言处理中顶尖模型在分类任务下共享一种通用的表征几何结构。类内变异并非随机,其分类相关成分与本类中心及竞争类中心存在强而有序的相关性。基于此,我们推导出一种主要由真实类与竞争类中心方向上的变异性以及全局归一化类半径驱动的解析均场理论,该理论补偿了真实表示中非高斯统计特性。该理论能准确预测不同架构和模态下的分类精度。相关几何量随模型规模系统性提升,与实际精度增益一致。理论的关键特征是稀疏性:仅需少量与真类及其最强对手相关的中心坐标即可实现精确预测,这使其与稀疏自编码器等方法相连接。这些结果为神经表示提供了简洁的预测理论,并表明深度网络中的分类受嵌于高维空间的稀疏、中心对齐结构支配。

原文摘要 · Abstract (English)

Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensional structure makes them difficult to interpret. We show that classification tasks give rise to a universal representational geometry, shared across state-of-the-art models in vision, audio, and language processing. The key structure is that within-class variability is not random in representation space. Instead, its classifier-relevant component has strong and structured correlations with the class's own centroid and with the centroids of its competing classes. Building on this observation, we derive an analytical mean-field theory governed mainly by the variability along true-class and rival-class centroid coordinates, together with a global renormalization of the class radius that compensates for the non-Gaussian statistics of real representations. The theory accurately predicts classification accuracy across architectures and modalities. The relevant geometric quantities improve systematically with model scale, mirroring the observed gains in accuracy. A striking feature of the theory is its sparsity: accurate prediction requires only a small set of centroid coordinates associated with the true class and its strongest rivals - connecting our framework to sparse-feature extraction approaches such as sparse autoencoders. Together, these results provide a parsimonious predictive theory of neural representations and suggest that classification in deep networks is governed by a sparse, centroid-aligned structure embedded within the full high-dimensional representation space.

神经表示分类几何稀疏性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。