arXiv:2608.27540stat.MLcs.IT2026-08

建立神经网络叠加的数学理论,揭示特征恢复的条件与边界。

Towards a mathematical theory of superposition

论文配图:Towards a mathematical theory of superposition
图 1 · 摘自论文原文
  • 用框架理论和压缩感知建模神经网络中稀疏特征的叠加与恢复。
  • 在随机支持下,可恢复期望稀疏度达 d/log n 的特征。
  • 首次给出实等角紧框架下的精确恢复阈值,对帧理论有独立价值。

我们利用框架理论和压缩感知工具,构建了神经网络中叠加现象的数学理论。模型中,稀疏二值向量 $x$ 通过过完备字典 $W$ 编码,通过 $ ext{ReLU}(W^ op W x + b)$ 与合适偏置 $b$ 实现特征恢复。证明了若干恢复定理:在随机支持设定下,对几乎紧致、低相干字典,当期望稀疏度达到 $d/ ext{log} n$ 量级时,可实现高概率支持恢复;在最坏情况支持下,给出了可计算的恢复条件,适用于高斯随机矩阵与等角紧框架。对于满足 $n > d+1$ 的实等角紧框架,基于格拉姆矩阵符号分布的新表征,确定了精确恢复阈值。该结果依赖于一个对框架理论具有独立意义的新刻画。

原文摘要 · Abstract (English)

We develop a mathematical theory of superposition in neural networks using tools from frame theory and compressed sensing. In our model, a sparse binary vector \(x\) of active features is encoded through an overcomplete dictionary \(W\), and feature recovery is performed by applying \(\operatorname{ReLU}(W^\top W x+b)\) with an appropriate bias vector \(b\). We prove several recovery theorems for this model. In the random-support setting, we establish high-probability support recovery for nearly tight, low-coherence dictionaries, with guarantees when the expected sparsity is up to order \(d/\log n\). In the worst-case support setting, we give a sharp and computable criterion for which sparsity levels permit support recovery. We apply this criterion to Gaussian random matrices and equiangular tight frames. For real equiangular tight frames with \(n>d+1\), we determine the exact recovery threshold in terms of the coherence. The proof of this result for real equiangular tight frames relies on a novel characterization---which should be of independent interest to frame theorists---of the distribution of signs in the Gram matrix.

神经网络压缩感知框架理论特征恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。