arXiv:2412.16738cs.LGcs.NA2024-12被引 28

提出新型神经网络KKAN,提升函数逼近与物理学习性能。

KKANs: Kurkova-Kolmogorov-Arnold Networks and Their Learning Dynamics

  • 双模块结构:内层用MLP,外层用基函数线性组合。
  • 在函数拟合和算子学习中优于MLP与原版KAN,接近最优MLP表现。
  • 揭示学习三阶段规律,动态注意力保障收敛稳定性。

受柯尔莫哥洛夫-阿诺德表示定理与库尔科娃近似原理启发,我们提出柯尔莫哥洛夫-库尔科娃-阿诺德网络(KKAN),一种包含两部分的新型架构:内层使用基于多层感知机(MLP)的强健函数,外层采用基函数的灵活线性组合。我们首先证明了KKAN是通用逼近器,并在函数回归、物理信息机器学习(PIML)和算子学习框架中展示了其广泛适用性。基准测试表明,KKAN在函数逼近和算子学习任务中优于MLP和原始柯尔莫哥洛夫-阿诺德网络(KAN),且在PIML任务中达到与全优化MLP相当的性能。为深入理解新模型行为,我们基于信息瓶颈理论分析其几何复杂度与学习动态,发现所有架构均呈现三种普遍学习阶段:拟合、过渡与扩散。我们发现几何复杂度与信噪比(SNR)存在强相关性,最优泛化出现在扩散阶段。此外,我们提出自缩放残差注意力权重,实现动态高SNR,确保均匀收敛并延长学习过程。

原文摘要 · Abstract (English)

Inspired by the Kolmogorov-Arnold representation theorem and Kurkova's principle of using approximate representations, we propose the Kurkova-Kolmogorov-Arnold Network (KKAN), a new two-block architecture that combines robust multi-layer perceptron (MLP) based inner functions with flexible linear combinations of basis functions as outer functions. We first prove that KKAN is a universal approximator, and then we demonstrate its versatility across scientific machine-learning applications, including function regression, physics-informed machine learning (PIML), and operator-learning frameworks. The benchmark results show that KKANs outperform MLPs and the original Kolmogorov-Arnold Networks (KANs) in function approximation and operator learning tasks and achieve performance comparable to fully optimized MLPs for PIML. To better understand the behavior of the new representation models, we analyze their geometric complexity and learning dynamics using information bottleneck theory, identifying three universal learning stages, fitting, transition, and diffusion, across all types of architectures. We find a strong correlation between geometric complexity and signal-to-noise ratio (SNR), with optimal generalization achieved during the diffusion stage. Additionally, we propose self-scaled residual-based attention weights to maintain high SNR dynamically, ensuring uniform convergence and prolonged learning.

神经网络函数逼近学习动态算子学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。