小模型靠结构先验,高效求解偏微分方程
Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE Solvers

- 用小波变换与多尺度特征金字塔构建物理先验
- 1000万参数模型媲美百万级大模型,波与声学问题提升显著
- 适合追求效率的物理模拟研究者使用
神经偏微分方程求解器沿用了视觉和语言模型的规模扩展路径,近期基础模型已达数十亿参数。我们主张:在该领域,规模无法替代架构先验;结构化先验能带来巨大参数效率,并且其成功与失败模式本身揭示了所捕获的物理特性。本文提出WaveLiT架构,结合离散小波变换实现无损多分辨率标记化、增强线性注意力块、共享权重多尺度特征金字塔及小波域辅助损失。针对八个TheWell基准测试,1-1000万参数的WaveLiT模型在性能上可比肩规模为自身100-1000倍的基础模型,尤其在波和声学主导的问题上表现最优,因小波-多尺度先验契合主要动力学结构,且单步误差在滚动计算中不几何累积。联合训练的1000万参数基础版本展现出结构化、可物理解释的迁移模式——在匹配先验的动力学场景下表现最强,而在混沌对流主导场景中最弱。整个流程可在单个GPU上完成训练。结果表明,小模型的性能主要由架构先验决定,而非规模;先验的失败模式是其内容的有效经验信号。
原文摘要 · Abstract (English)
Neural PDE solvers have followed the scaling trajectory of vision and language, with recent foundation models reaching billions of parameters. We argue that scale is a poor substitute for architectural inductive bias in this domain: structured priors deliver outsized parameter efficiency, and the pattern of where they succeed and fail is itself informative about what they capture. We instantiate this argument in WaveLiT, an architecture combining a discrete wavelet transform for lossless multi-resolution tokenization, an augmented linear attention block, a shared-weight multiscale feature pyramid, and a wavelet-domain auxiliary loss. Bespoke 1-10M-parameter WaveLiT models compete with foundation models of 100-1000$\times$ their size across eight TheWell benchmarks, with the largest gains on wave and acoustic-dominated benchmarks where the wavelet-multiscale prior fits the dominant dynamical structure and small per-step errors do not compound geometrically under rollout. Trained jointly across all eight benchmarks, a 10M-parameter foundation variant exhibits a structured, physically interpretable transfer pattern -- strongest where the wavelet-multiscale prior matches the dynamics, weakest on chaotic advection-dominated flows. The entire pipeline trains on a single GPU. The results suggest that small-model PDE performance is shaped by architectural inductive bias rather than scale, and that the structure of a prior's failures is a useful empirical signal about its content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。