arXiv:2411.05296cs.LGcs.AI2024-11被引 7

KAN网络在高维数据上表现优于MLP,但训练更不稳定。

On Training of Kolmogorov-Arnold Networks

  • 采用不同初始化、优化器和学习率对比KAN与MLP训练动态
  • 在高维数据集上测试准确率显示KAN参数效率更高
  • 提出提升大模型训练稳定性的实用建议

Kolmogorov-Arnold Networks(KAN)最近被提出作为多层感知机(MLP)架构的灵活替代方案。本文研究了不同KAN架构的训练动态,并与对应MLP模型进行比较。实验涵盖多种初始化策略、优化器和学习率设置,还采用了无需反向传播的HSIC瓶颈方法。结果表明,在高维数据集上,以测试准确率为标准,KAN是MLP的有效替代方案,且具有更高的参数效率,但训练过程更不稳定。最后,本文针对大规模KAN模型的训练稳定性问题提出了改进建议。

原文摘要 · Abstract (English)

Kolmogorov-Arnold Networks have recently been introduced as a flexible alternative to multi-layer Perceptron architectures. In this paper, we examine the training dynamics of different KAN architectures and compare them with corresponding MLP formulations. We train with a variety of different initialization schemes, optimizers, and learning rates, as well as utilize back propagation free approaches like the HSIC Bottleneck. We find that (when judged by test accuracy) KANs are an effective alternative to MLP architectures on high-dimensional datasets and have somewhat better parameter efficiency, but suffer from more unstable training dynamics. Finally, we provide recommendations for improving training stability of larger KAN models.

神经网络深度学习训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。