用分数阶扩散模拟多尺度注意力,提升Transformer效率与表达力。
Fractional neural attention for efficient multiscale sequence processing
- 基于分数拉普拉斯算子建模分形扩散过程,实现多尺度依赖统一建模。
- 单层单头即达竞争力文本分类性能,图像与翻译任务亦有提升。
- 兼具生物可解释性与计算高效性,适合关注模型机制与效率的研究者。
注意力机制是Transformer模型计算能力的核心,但其自注意力原理的理解与拓展仍是人工智能发展的关键挑战。受生物注意机制的多尺度动态和动力系统理论启发,本文提出分数阶神经注意力(Fractional Neural Attention, FNA),一种源于神经科学的多尺度信息处理原则性框架。FNA通过由分数拉普拉斯算子控制的莱维扩散建模标记间交互,天然实现跨多尺度的短程与长程依赖。该机制提升了表达能力并加速信息混合,增强Transformer的基础计算能力。理论上,FNA的动力学由分数扩散方程支配,其注意力网络表现出更大的谱间隙和更短路径长度——计算效率增强的机制标志。实验上,即使仅用单层单头,FNA在文本分类中也达到竞争力表现;在图像处理和神经机器翻译任务中亦有性能提升。此外,来自几何调和分析的扩散图算法可在保留嵌入与隐藏状态内在结构的前提下,对FNA权重进行降维。这些结果共同确立了FNA作为连接自注意力、随机动力学与几何结构的原理性机制,为强大且具生物学依据的人工智能提供可解释的理论基础。
原文摘要 · Abstract (English)
Attention mechanisms underpin the computational power of Transformer models, which have achieved remarkable success across diverse domains. Yet understanding and extending the principles underlying self-attention remains a key challenge for advancing artificial intelligence. Drawing inspiration from the multiscale dynamics of biological attention and from dynamical systems theory, we introduce Fractional Neural Attention (FNA), a principled, neuroscience-inspired framework for multiscale information processing. FNA models token interactions through Lévy diffusion governed by the fractional Laplacian, intrinsically realizing simultaneous short- and long-range dependencies across multiple scales. This mechanism yields greater expressivity and faster information mixing, advancing the foundational capacity of Transformers. Theoretically, we show that FNA's dynamics are governed by the fractional diffusion equation, and that the resulting attention networks exhibit larger spectral gaps and shorter path lengths -- mechanistic signatures of enhanced computational efficiency. Empirically, FNA achieves competitive text-classification performance even with a single layer and a single head; it also improves performance in image processing and neural machine translation. Finally, the diffusion map algorithm from geometric harmonics enables dimensionality reduction of FNA weights while preserving the intrinsic structure of embeddings and hidden states. Together, these results establish FNA as a principled mechanism connecting self-attention, stochastic dynamics, and geometry, providing an interpretable, biologically grounded foundation for powerful, neuroscience-inspired AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。