arXiv:2605.08856cs.LG2026-05

通过抑制误差瞬态放大,显著提升神经物理模拟器的长期预测能力

Controlling Transient Amplification Improves Long-horizon Rollouts

论文配图:Controlling Transient Amplification Improves Long-horizon Rollouts
图 1 · 摘自论文原文
  • 提出对偶性正则化,降低雅可比矩阵非正规与不可交换性
  • 在数千步滚动预测中保持精度,即使系统整体稳定仍有效
  • 适用于需要长时序模拟的科学计算与气候预测场景

自回归神经模拟器在短时物理系统预测上已接近经典求解器,但在长时滚动预测中精度迅速下降。本文识别出轨迹附近扰动的瞬态放大是导致滚动误差的关键结构机制。通过线性化分析发现,当自回归轨迹上的雅可比矩阵非正规且不可交换时,模型会瞬时放大误差,引发滚动漂移,即使整体系统渐近稳定亦然。基于此,提出对偶性正则化:结合两项惩罚,分别减小单步雅可比矩阵的非正规性缺陷与跨步雅可比矩阵的对易子范数。惩罚项通过雅可比-向量乘积估计,无推理开销。我们推导了近似对偶性与正规性下的传播器界,量化了滚动误差。在1D与2D时空数据的合成及真实场景中,评估UNet与FNO变体,均实现数千步成功滚动预测。进一步表明,该方法在不使用新数据的前提下提升了FourCastNet对ERA5气候预报的性能,尤其在分布外情形下优势明显:训练于数百步轨迹,正则化模型在基线发散的初始条件下仍能保持分布内长达数千步的滚动。

原文摘要 · Abstract (English)

Autoregressive neural simulators now match classical solvers on short-horizon prediction of physical systems, yet their accuracy degrades rapidly when rolled out over long horizons. In this work, we identify transient amplification of perturbations around rollout trajectories as a structural mechanism driving rollout error. Using a linearization analysis we show that when the Jacobians along an autoregressive trajectory are non-normal and non-commuting, the model amplifies errors transiently, resulting in model rollout drift even when the overall system is asymptotically stable. Building on the analysis, we propose commutativity regularization: a combination of two penalties designed to reduce the normality defect of individual Jacobians and the commutator norm of Jacobians across steps. The penalties are estimated with Jacobian-vector products and have no inference-time cost. We show a propagator bound that quantifies rollout error under approximate commutativity and normality. We evaluate UNet and FNO variants with commutativity regularization on 1D and 2D spatio-temporal data in synthetic and real settings, showing successful long-horizon rollouts over thousands of steps. Further, we show that the method improves FourCastNet climate forecasts on ERA5 without using any new data. The gain is most pronounced out-of-distribution: trained on trajectories of a few hundred steps, regularized models remain in-distribution for thousands of rollout steps on initial conditions where baselines diverge.

物理模拟长时预测稳定性优化神经算子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。