arXiv:2602.02224cs.LGcs.AI2026-02被引 2

用谱分析揭示神经网络特征的几何结构,解释为何超维度表示会局部化。

Spectral Superposition: A Theory of Feature Geometry

  • 通过权重矩阵的谱特性研究特征在空间中的分布形态。
  • 发现容量饱和导致特征集中于单一特征空间,形成离散分类结构。
  • 适用于真实模型诊断,为可解释性提供新理论框架。

神经网络在维度少于特征数的情况下,通过超叠加机制共享表征空间。现有方法虽能分解激活为稀疏线性特征,却忽略了几何结构。本文提出基于权重矩阵谱(如特征值、特征空间)的理论体系,引入框架算子 $F = WW^ op$,以谱测度描述每个特征如何分配范数至各特征空间。相比以往仅刻画特征对间交互的方法,该框架能捕捉所有特征的全局几何关系。在超叠加的简化模型中,证明容量饱和将迫使特征坍缩至单一特征空间,组织成紧框架,并可通过关联方案实现离散分类,涵盖先前所有几何构型(如单形、多边形、反棱柱)。该谱测度形式适用于任意权重矩阵,可在非玩具设置中诊断特征局域化现象。这些结果指向一个更广泛的计划:将算子理论应用于模型可解释性。

原文摘要 · Abstract (English)

Neural networks represent more features than they have dimensions via superposition, forcing features to share representational space. Current methods decompose activations into sparse linear features but discard geometric structure. We develop a theory for studying the geometric structre of features by analyzing the spectra (eigenvalues, eigenspaces, etc.) of weight derived matrices. In particular, we introduce the frame operator $F = WW^\top$, which gives us a spectral measure that describes how each feature allocates norm across eigenspaces. While previous tools could describe the pairwise interactions between features, spectral methods capture the global geometry (``how do all features interact?''). In toy models of superposition, we use this theory to prove that capacity saturation forces spectral localization: features collapse onto single eigenspaces, organize into tight frames, and admit discrete classification via association schemes, classifying all geometries from prior work (simplices, polygons, antiprisms). The spectral measure formalism applies to arbitrary weight matrices, enabling diagnosis of feature localization beyond toy settings. These results point toward a broader program: applying operator theory to interpretability.

神经网络特征几何谱分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。