预训练分词器可显著提升物理模拟模型的效率与精度。
On the Value of Tokeniser Pretraining in Physics Foundation Models
- 先用自编码任务预训练分词器,再训练动力学模型。
- 同领域预训练使10500步后VRMSE降低64%。
- 适合需高效物理模拟的科研与工程场景。
我们研究了分词器预训练对物理模拟精度与效率的影响。现代高分辨率模拟产生大量跨越不同物理状态与尺度的数据。训练基础模型以学习底层动力学,有助于在数据稀缺情况下建模复杂多物理现象。现有物理基础模型通常联合学习两项任务:(i) 高分辨率时空数据的紧凑表示,(ii) 控制性物理动力学。但同时从零开始学习两者会削弱效果。我们发现,先用自编码目标预训练分词器,再训练动力学模型,能显著提升计算效率。值得注意的是,该收益大小取决于领域一致性:在相同物理系统上预训练带来最大提升,其他系统则有中等收益。同领域预训练在10,500次训练步骤后使VRMSE降低64%。据我们所知,这是首个针对物理基础模型分词器预训练的系统性研究。我们还引入灵活的时空压缩操作,扩展因果卷积支持运行时可调压缩比,实现对下游任务的高效适配。研究为训练高效物理模拟器提供了实用指导,并强调了预训练数据选择的战略意义。
原文摘要 · Abstract (English)
We investigate the impact of tokeniser pretraining on the accuracy and efficiency of physics emulation. Modern high-resolution simulations produce vast volumes of data spanning diverse physical regimes and scales. Training foundation models to learn the dynamics underlying such data enables the modelling of complex multiphysics phenomena, especially in data-limited settings. The emerging class of physics foundation models typically aims to learn two tasks jointly: (i) extracting compact representations of high-resolution spatiotemporal data, and (ii) capturing governing physical dynamics. However, learning both tasks from scratch simultaneously can impede the effectiveness of either process. We show that pretraining the tokeniser with an autoencoding objective prior to training the dynamics model enhances computational efficiency for physics emulation. Notably, the magnitude of this benefit depends on domain alignment: pretraining on the same physical system as the emulation task yields the largest improvements, while pretraining on other systems provides moderate gains. In-domain pretraining reduces VRMSE by 64% after 10,500 training steps compared to training from scratch. To our knowledge, this is the first systematic investigation of tokeniser pretraining for physics foundation models. We further introduce flexible spatiotemporal compression operations that extend causal convolutions to support runtime-adjustable compression ratios, enabling efficient adaptation to diverse downstream tasks. Our findings provide practical guidance for training efficient physics emulators and highlight the importance of strategic pretraining data selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。