arXiv:2503.23038cs.LG2025-03被引 1

用核方法统一KAN与自注意力,降低50%参数量

Function Fitting Based on Kolmogorov-Arnold Theorem and Kernel Functions

  • 基于科尔莫戈罗夫-阿诺德定理与核函数构建统一框架
  • 提出低秩伪多头注意力,参数减少近50%
  • 验证非线性核在特征提取中的有效性,适合视觉建模

本文提出一种基于科尔莫戈罗夫-阿诺德表示定理与核方法的统一理论框架。通过分析核函数、KAN中的B样条基函数及自注意力机制中的内积运算之间的数学关系,建立了以核函数线性组合为核心的特征拟合框架。在此基础上,提出低秩伪多头自注意力模块(Pseudo-MHSA),将传统MHSA的参数量减少近50%。进一步设计高斯核多头自注意力变体(Gaussian-MHSA),验证了非线性核函数在特征提取中的有效性。在CIFAR-10数据集上的实验表明,Pseudo-MHSA模型在MAE框架下性能与同维度ViT相当,可视化分析显示两者多头分布模式高度相似。代码已公开。

原文摘要 · Abstract (English)

This paper proposes a unified theoretical framework based on the Kolmogorov-Arnold representation theorem and kernel methods. By analyzing the mathematical relationship among kernels, B-spline basis functions in Kolmogorov-Arnold Networks (KANs) and the inner product operation in self-attention mechanisms, we establish a kernel-based feature fitting framework that unifies the two models as linear combinations of kernel functions. Under this framework, we propose a low-rank Pseudo-Multi-Head Self-Attention module (Pseudo-MHSA), which reduces the parameter count of traditional MHSA by nearly 50\%. Furthermore, we design a Gaussian kernel multi-head self-attention variant (Gaussian-MHSA) to validate the effectiveness of nonlinear kernel functions in feature extraction. Experiments on the CIFAR-10 dataset demonstrate that Pseudo-MHSA model achieves performance comparable to the ViT model of the same dimensionality under the MAE framework and visualization analysis reveals their similarity of multi-head distribution patterns. Our code is publicly available.

核方法自注意力模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。