arXiv:2505.18131cs.LGcs.AI2025-05被引 6

用KAN结构加速多通道MLP训练,提升精度与速度

Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement

  • 基于KAN的几何局部支撑特性,构建预条件优化框架
  • 训练速度显著提升,精度优于传统MLP,跨任务验证有效
  • 适合科学计算与回归任务,对模型加速研究者有参考价值

多层感知机(MLPs)是现代深度学习中的核心架构。近年来,柯尔莫哥洛夫-阿诺德网络(KANs)因其在科学机器学习等任务中的出色表现而受到关注。本文揭示了KAN与多通道MLP之间的结构等价关系,发现其基函数具有几何局部支持性,并可在ReLU基下实现预条件梯度下降,从而加速训练并提升精度。我们证明自由节点样条KAN架构等价于一类沿权重张量通道维度进行几何精炼的MLP。基于此,提出层级精炼方案,大幅加速多通道MLP训练。进一步通过联合优化样条节点的一维位置与权重,可进一步提升性能。该方法在多个回归与科学机器学习基准上得到验证。

原文摘要 · Abstract (English)

Multilayer perceptrons (MLPs) are a workhorse machine learning architecture, used in a variety of modern deep learning frameworks. However, recently Kolmogorov-Arnold Networks (KANs) have become increasingly popular due to their success on a range of problems, particularly for scientific machine learning tasks. In this paper, we exploit the relationship between KANs and multichannel MLPs to gain structural insight into how to train MLPs faster. We demonstrate the KAN basis (1) provides geometric localized support, and (2) acts as a preconditioned descent in the ReLU basis, overall resulting in expedited training and improved accuracy. Our results show the equivalence between free-knot spline KAN architectures, and a class of MLPs that are refined geometrically along the channel dimension of each weight tensor. We exploit this structural equivalence to define a hierarchical refinement scheme that dramatically accelerates training of the multi-channel MLP architecture. We show further accuracy improvements can be had by allowing the $1$D locations of the spline knots to be trained simultaneously with the weights. These advances are demonstrated on a range of benchmark examples for regression and scientific machine learning.

KANMLP加速科学计算几何精炼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。