MATEY动态调整图像块大小,提升物理系统时空建模效率与精度
MATEY: multiscale adaptive foundation models for spatiotemporal physical systems
- 根据局部特征自适应调整图像块大小,减少冗余计算
- 自适应分块使精度提升,且令牌序列长度几乎不变
- 预训练于PDEBench数据在少样本下表现更优,适合低数据场景
使用视觉变换器(ViT)架构精确表征时空物理系统的多尺度特征,需要极长且计算成本高昂的令牌序列。为此,我们提出两种自适应分块方案,根据局部特征动态调整图像块大小:一种确保收敛到均匀细化,另一种更具计算效率。此外,我们设计了一系列时空注意力机制,将时间或轴向空间维度解耦,并评估其计算与数据效率。通过一系列实验评估所提出的多尺度自适应模型MATEY。结果表明,自适应分块方案在不显著增加令牌序列长度的前提下提升了精度。相比全时空注意力或仅解耦时间维度的方案,完全解耦轴向注意力效率较低、表达能力较弱,需更多训练时间与模型参数才能达到相同精度。最后,我们在两个具有不同物理特性的微调任务中证明,基于PDEBench数据预训练的模型优于从零开始训练的模型,尤其在注意力层冻结的低数据情形下表现更佳。
原文摘要 · Abstract (English)
Accurate representation of the multiscale features in spatiotemporal physical systems using vision transformer (ViT) architectures requires extremely long, computationally prohibitive token sequences. To address this issue, we propose two adaptive tokenization schemes that dynamically adjust patch sizes based on local features: one ensures convergent behavior to uniform patch refinement, while the other offers better computational efficiency. Moreover, we present a set of spatiotemporal attention schemes, where the temporal or axial spatial dimensions are decoupled, and evaluate their computational and data efficiencies. We assess the performance of the proposed multiscale adaptive model, MATEY, in a sequence of experiments. The results show that adaptive tokenization schemes achieve improved accuracy without significantly increasing the length of the token sequence. Compared to a full spatiotemporal attention scheme or a scheme that decouples only the temporal dimension, we find that fully decoupled axial attention is less efficient and expressive, requiring more training time and model weights to achieve the same accuracy. Finally, we demonstrate in two fine-tuning tasks featuring different physics that models pretrained on PDEBench data outperform the ones trained from scratch, especially in the low data regime with frozen attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。