arXiv:2509.15536cs.CVcs.RO2025-09NeurIPS被引 3

SAMPO通过分尺度自回归与运动提示,提升视频生成的连贯性与效率。

SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models

  • 分尺度自回归结合因果解码,支持并行生成且保持空间结构
  • 生成质量显著提升,推理速度加快4.4倍
  • 适合需要高效动态场景模拟的强化学习与规划任务

世界模型使智能体能在想象环境中模拟动作后果,用于规划、控制和长时决策。然而,现有自回归世界模型在视觉连贯性上受限于空间结构破坏、解码效率低和运动建模不足。为此,我们提出规模分层自回归与运动提示框架(SAMPO),融合帧内自回归生成与帧间因果建模。SAMPO采用时间因果解码与双向空间注意力,保持空间局部性并支持各尺度内并行解码,显著提升时序一致性与推演效率。设计非对称多尺度分词器,在保留观测帧空间细节的同时,为未来帧提取紧凑动态表示,优化内存与性能。引入轨迹感知运动提示模块,注入物体与机器人轨迹的时空线索,聚焦动态区域,增强时序一致性和物理真实性。大量实验表明,SAMPO在动作条件视频预测和基于模型的控制中表现优异,生成质量更高,推理速度提升4.4倍。还评估了零样本泛化与可扩展性,证明其能泛化到未见任务,并随模型增大而持续受益。

原文摘要 · Abstract (English)

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and inadequate motion modeling. In response, we propose \textbf{S}cale-wise \textbf{A}utoregression with \textbf{M}otion \textbf{P}r\textbf{O}mpt (\textbf{SAMPO}), a hybrid framework that combines visual autoregressive modeling for intra-frame generation with causal modeling for next-frame generation. Specifically, SAMPO integrates temporal causal decoding with bidirectional spatial attention, which preserves spatial locality and supports parallel decoding within each scale. This design significantly enhances both temporal consistency and rollout efficiency. To further improve dynamic scene understanding, we devise an asymmetric multi-scale tokenizer that preserves spatial details in observed frames and extracts compact dynamic representations for future frames, optimizing both memory usage and model performance. Additionally, we introduce a trajectory-aware motion prompt module that injects spatiotemporal cues about object and robot trajectories, focusing attention on dynamic regions and improving temporal consistency and physical realism. Extensive experiments show that SAMPO achieves competitive performance in action-conditioned video prediction and model-based control, improving generation quality with 4.4$\times$ faster inference. We also evaluate SAMPO's zero-shot generalization and scaling behavior, demonstrating its ability to generalize to unseen tasks and benefit from larger model sizes.

世界模型视频生成自回归运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。