arXiv:2505.12252cs.LG2025-05被引 1

用多项式基重构注意力,提速且不丢精度。

SchoenbAt: Rethinking Attention with Polynomial basis

  • 基于Schoenberg定理,用多项式基替代傅里叶基近似注意力。
  • 在多个数据集上实现更快计算速度,精度仍保持竞争力。
  • 适合需要高效注意力的模型部署场景,如长序列处理。

核化注意力通过核函数建模序列相关性,显著优化了注意力机制。在调和分析理论保证下,核函数可展开为基函数,催生基于随机特征的方法以提升核化注意力效率并保持预测性能。然而,现有方法仅限于Bochner定理下的傅里叶基展开。本文提出基于Schoenberg定理的注意力(SchoenbAt),通过随机Maclaurin特征,以多项式基近似点积核化注意力,并采用两阶段正则化约束输入空间与恢复输出尺度,可作为点积核化注意力的即插即用替代方案。理论上证明了SchoenbAt的无偏性及浓度误差界,支持其高效与准确;实证验证了其在多种随机特征维度下的有效性。真实数据集评估显示,SchoenbAt显著提升计算速度,同时保持高精度,优于多个高效注意力方法。

原文摘要 · Abstract (English)

Kernelized attention extends the attention mechanism by modeling sequence correlations through kernel functions, making significant progresses in optimizing attention. Under the guarantee of harmonic analysis theory, kernel functions can be expanded with basis functions, inspiring random feature-based approaches to enhance the efficiency of kernelized attention while maintaining predictive performance. However, current random feature-based works are limited to the Fourier basis expansions under Bochner's theorem. We propose Schoenberg's theorem-based attention (SchoenbAt), which approximates dot-product kernelized attention with the polynomial basis under Schoenberg's theorem via random Maclaurin features and applies a two-stage regularization to constrain the input space and restore the output scale, acting as a drop-in replacement of dot-product kernelized attention. Our theoretical proof of the unbiasedness and concentration error bound of SchoenbAt supports its efficiency and accuracy as a kernelized attention approximation, which is also empirically validated under various random feature dimensions. Evaluations on real-world datasets demonstrate that SchoenbAt significantly enhances computational speed while preserving competitive performance in terms of precision, outperforming several efficient attention methods.

注意力机制核方法多项式基高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。