让大模型零成本复用小模型调参结果,大幅降低复杂方程求解的计算开销。
Maximal Update Parametrization and Zero-Shot Hyperparameter Transfer for Fourier Neural Operators
- 基于μP框架设计参数化方案,实现不同规模傅里叶神经算子间的参数迁移
- 在多种偏微分方程上验证,大模型直接使用小模型最优超参数仍保持高精度
- 无需重新调参即可部署千亿参数模型,适合大规模科学计算场景
傅里叶神经算子(FNO)为求解复杂偏微分方程提供了系统性方法。然而,为应对更复杂的方程,需增加傅里叶模式数量,这显著增加了模型参数量,使超参数调优计算成本过高。为此,我们提出μTransfer-FNO,一种零样本超参数迁移技术,可将小规模FNO上训练出的最优配置直接应用于百亿参数级大模型,无需额外调参。基于最大更新参数化(μP)框架,我们数学推导出一种参数化方案,支持在不同傅里叶模式数的FNO间迁移最优超参数,并通过多种偏微分方程的大量实验加以验证。实证表明,Transfer-FNO显著降低了大规模FNO的调参计算成本,同时保持或提升模型精度。
原文摘要 · Abstract (English)
Fourier Neural Operators (FNOs) offer a principled approach for solving complex partial differential equations (PDEs). However, scaling them to handle more complex PDEs requires increasing the number of Fourier modes, which significantly expands the number of model parameters and makes hyperparameter tuning computationally impractical. To address this, we introduce $μ$Transfer-FNO, a zero-shot hyperparameter transfer technique that enables optimal configurations, tuned on smaller FNOs, to be directly applied to billion-parameter FNOs without additional tuning. Building on the Maximal Update Parametrization ($μ$P) framework, we mathematically derive a parametrization scheme that facilitates the transfer of optimal hyperparameters across models with different numbers of Fourier modes in FNOs, which is validated through extensive experiments on various PDEs. Our empirical study shows that Transfer-FNO reduces computational cost for tuning hyperparameters on large FNOs while maintaining or improving accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。