用可学习激活函数提升物理信息神经网络求解微分方程的精度与稳定性。
Learnable Activation Functions in Physics-Informed Neural Networks for Solving Partial Differential Equations
- 引入可学习激活函数,缓解传统PINN的频谱偏差和收敛不稳问题。
- 实验表明低频偏差并非唯一关键,更宽的NTK谱可能导致收敛失败。
- 不同微分方程类型需匹配特定激活形式,设计需因题而异。
物理信息神经网络(PINNs)已成为求解偏微分方程(PDEs)的有前景方法,但存在频谱偏差(倾向于学习低频成分,难以捕捉高频特征)和收敛动力学不稳定(主要源于多目标损失函数)等问题,影响其对快速振荡、陡峭梯度和复杂边界行为问题的求解精度。本文系统研究可学习激活函数作为解决方案,比较采用固定与可学习激活函数的多层感知机(MLPs)及使用可学习基函数的柯尔莫哥洛夫-阿诺德网络(KANs)。评估涵盖线性与非线性波问题、多物理系统和流体动力学等多种PDE类型。通过经验神经正切核(NTK)分析与海森矩阵特征值分解,评估模型的频谱偏差与收敛稳定性。结果揭示表达能力与训练稳定性之间存在权衡:可学习激活函数在简单架构中表现良好,但在复杂网络中因函数维度更高而出现可扩展性问题。反直觉的是,仅降低频谱偏差不足以保证更高精度,因为具有更广NTK特征值谱的函数可能表现出收敛不稳定性。我们证明激活函数选择本质上依赖于具体问题,不同基函数对特定PDE特性具有独特优势。这些发现将有助于设计更鲁棒的神经PDE求解器。
原文摘要 · Abstract (English)
Physics-Informed Neural Networks (PINNs) have emerged as a promising approach for solving Partial Differential Equations (PDEs). However, they face challenges related to spectral bias (the tendency to learn low-frequency components while struggling with high-frequency features) and unstable convergence dynamics (mainly stemming from the multi-objective nature of the PINN loss function). These limitations impact their accuracy for problems involving rapid oscillations, sharp gradients, and complex boundary behaviors. We systematically investigate learnable activation functions as a solution to these challenges, comparing Multilayer Perceptrons (MLPs) using fixed and learnable activation functions against Kolmogorov-Arnold Networks (KANs) that employ learnable basis functions. Our evaluation spans diverse PDE types, including linear and non-linear wave problems, mixed-physics systems, and fluid dynamics. Using empirical Neural Tangent Kernel (NTK) analysis and Hessian eigenvalue decomposition, we assess spectral bias and convergence stability of the models. Our results reveal a trade-off between expressivity and training convergence stability. While learnable activation functions work well in simpler architectures, they encounter scalability issues in complex networks due to the higher functional dimensionality. Counterintuitively, we find that low spectral bias alone does not guarantee better accuracy, as functions with broader NTK eigenvalue spectra may exhibit convergence instability. We demonstrate that activation function selection remains inherently problem-specific, with different bases showing distinct advantages for particular PDE characteristics. We believe these insights will help in the design of more robust neural PDE solvers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。