用新方法让自回归模型稳定预测混沌系统,参数少65倍还更快。
A Hybridizable Neural Time Integrator for Stable Autoregressive Forecasting

- 把Transformer嵌入新型有限元框架,利用拓扑结构保证稳定性。
- 训练时梯度有界,避免爆炸;推理时离散能量守恒,长期预测不发散。
- 适合做小规模高精度科学模拟,尤其适合计算资源有限的场景。
针对长时间跨度下混沌动力系统的自回归建模,训练与推理的稳定性是构建科学基础模型的核心挑战。本文提出一种混合方法:将自回归Transformer嵌入基于射击法的新型混合有限元方案中,揭示了可证明稳定的拓扑结构。对于正问题,证明了离散能量的保持;对于训练,证明了梯度的统一有界性,可有效避免梯度爆炸。结合视觉Transformer后,生成的隐空间标记具备结构保持的动力学特性。相比现代基础模型,本方法实现参数量减少65倍,并在长时序预测中表现更优。以聚变部件的“微型基础模型”为例,仅需12次仿真即可训练出实时代理模型,相比粒子-网格模拟实现9000倍加速。
原文摘要 · Abstract (English)
For autoregressive modeling of chaotic dynamical systems over long time horizons, the stability of both training and inference is a major challenge in building scientific foundation models. We present a hybrid technique in which an autoregressive transformer is embedded within a novel shooting-based mixed finite element scheme, exposing topological structure that enables provable stability. For forward problems, we prove preservation of discrete energies, while for training we prove uniform bounds on gradients, provably avoiding the exploding gradient problem. Combined with a vision transformer, this yields latent tokens admitting structure-preserving dynamics. We outperform modern foundation models with a $65\times$ reduction in model parameters and long-horizon forecasting of chaotic systems. A "mini-foundation" model of a fusion component shows that 12 simulations suffice to train a real-time surrogate, achieving a $9{,}000\times$ speedup over particle-in-cell simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。