arXiv:2604.15769cs.LGcs.AI2026-04

首次建立脉冲Transformer表达理论,解释其低功耗高效运行原理。

Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension

  • 基于漏电积分-放电神经元,构建可逼近任意排列等变函数的脉冲注意力机制
  • 推导出严格脉冲数下界:ε精度需Ω(L_f²nd/ε²)个脉冲,实验验证预测相关性R²=0.97
  • 提出输入依赖的有效维度概念,解释为何4个时间步即可满足性能需求

脉冲Transformer在类脑硬件上实现与传统Transformer相当的准确率,同时能耗降低38至57倍,但缺乏理论指导。本文首次建立脉冲自注意力的完整表达能力理论。证明使用漏电积分-放电神经元的脉冲注意力是连续排列等变函数的通用近似器,并给出新型侧抑制网络实现softmax归一化,收敛速度达O(1/√T)。通过速率-失真理论推导严格的脉冲数下界:ε-近似需Ω(L_f²nd/ε²)个脉冲,具信息论基础。关键洞察在于引入输入依赖的有效维度(d_eff=47–89,CIFAR/ImageNet),解释为何仅需T=4个时间步即可满足性能,尽管最坏情况需T≥10,000。提供带校准常数(C=2.3,95%置信区间[1.9, 2.7])的设计规则。在Spikformer、QKFormer和SpikingResformer上的视觉与语言基准实验验证预测,相关性R²=0.97(p<0.001)。本框架为类脑Transformer设计提供了首个原理性基础。

原文摘要 · Abstract (English)

Spiking transformers achieve competitive accuracy with conventional transformers while offering $38$-$57\times$ energy efficiency on neuromorphic hardware, yet no theoretical framework guides their design. This paper establishes the first comprehensive expressivity theory for spiking self-attention. We prove that spiking attention with Leaky Integrate-and-Fire neurons is a universal approximator of continuous permutation-equivariant functions, providing explicit spike circuit constructions including a novel lateral inhibition network for softmax normalization with proven $O(1/\sqrt{T})$ convergence. We derive tight spike-count lower bounds via rate-distortion theory: $\varepsilon$-approximation requires $Ω(L_f^2 nd/\varepsilon^2)$ spikes, with rigorous information-theoretic derivation. Our key insight is input-dependent bounds using measured effective dimensions ($d_{\text{eff}}=47$--$89$ for CIFAR/ImageNet), explaining why $T=4$ timesteps suffice despite worst-case $T \geq 10{,}000$ predictions. We provide concrete design rules with calibrated constants ($C=2.3$, 95\% CI: $[1.9, 2.7]$). Experiments on Spikformer, QKFormer, and SpikingResformer across vision and language benchmarks validate predictions with $R^2=0.97$ ($p<0.001$). Our framework provides the first principled foundation for neuromorphic transformer design.

脉冲神经网络类脑计算注意力机制理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。