arXiv:2510.20807cs.CVcs.LG2025-10被引 2

用纯Transformer实现更长时序的物理仿真视频预测

Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers

  • 基于像素空间的自回归Transformer,无需复杂训练策略
  • 相比现有方法,物理准确预测时间延长50%
  • 模型可解释性强,能推断出物理参数并泛化到新场景

受自回归大语言模型性能与可扩展性的启发,基于Transformer的模型在视觉领域取得成功。本文研究一种适用于视频预测的Transformer适配方法,采用端到端简单架构,比较多种时空自注意力布局。聚焦于对物理模拟随时间演化的因果建模,这是现有视频生成方法的常见短板;我们通过物理对象追踪指标和无监督训练物理模拟数据集,尝试分离时空推理能力。提出一种简单有效的纯Transformer模型用于自回归视频预测,利用连续像素空间表示进行预测。无需复杂训练策略或潜在特征学习组件,该方法在物理准确性预测上将时间范围扩展了最多50%,同时在常用视频质量指标上保持相当表现。此外,我们通过探测模型开展可解释性实验,发现网络部分区域编码了有助于准确估计偏微分方程(PDE)模拟参数的信息,且该能力可泛化至分布外模拟参数。本工作为基于注意力机制的时空建模提供了一个简洁、参数高效且可解释的平台。

原文摘要 · Abstract (English)

Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study investigates a transformer adaptation for video prediction with a simple end-to-end approach, comparing various spatiotemporal self-attention layouts. Focusing on causal modeling of physical simulations over time; a common shortcoming of existing video-generative approaches, we attempt to isolate spatiotemporal reasoning via physical object tracking metrics and unsupervised training on physical simulation datasets. We introduce a simple yet effective pure transformer model for autoregressive video prediction, utilizing continuous pixel-space representations for video prediction. Without the need for complex training strategies or latent feature-learning components, our approach significantly extends the time horizon for physically accurate predictions by up to 50% when compared with existing latent-space approaches, while maintaining comparable performance on common video quality metrics. In addition, we conduct interpretability experiments to identify network regions that encode information useful to perform accurate estimations of PDE simulation parameters via probing models, and find that this generalizes to the estimation of out-of-distribution simulation parameters. This work serves as a platform for further attention-based spatiotemporal modeling of videos via a simple, parameter efficient, and interpretable approach.

视频预测Transformer物理模拟可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。