arXiv:2608.26239cs.RO2026-08

让机器人模型能长期、可控地预测未来视觉,支持分钟级连续模拟。

WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression

论文配图:WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression
图 1 · 摘自论文原文
  • 分尺度自回归生成未来画面,动作与后果显式关联。
  • 在有限内存下实现分钟级连贯仿真,减少动作漂移和长期不一致。
  • 适合需要长时序规划的机器人学习与策略优化场景。

生成式世界模型为机器人提供交互下的世界演化预测能力,具备仿真、规划、策略评估和学习的巨大潜力。超越片段级未来预测,统一的生成范式应关联动作与后果,支持灵活时长和连续交互,并实现奖励驱动优化。我们提出WALL-SS,一种通过分尺度自回归扩展生成视觉未来的世界模型,实现动作可控的长时程机器人仿真。WALL-SS将具身轨迹表示为观察与动作的时间交错因果序列,明确动作依赖的状态转移,自然支持变长生成、可复用因果状态的流式扩展,以及通过序列概率直接优化。为在长时程下有效建模,我们以粗到细的方式生成每帧未来图像,并在同一层级构建三个互补组件:动作条件的下一级预测注入对齐尺度的动作表征,增强动作-未来耦合,建模成功与失败行为;分尺度压缩的长时记忆在细粒度保留近期交互的同时压缩远期观测与动作,尺度级梦境强迫提升对自生成上下文的鲁棒性;策略内对齐则通过动作跟随与长期一致性奖励优化自回归视觉动态,同时保持预训练视觉分布。实验表明,WALL-SS显著提升动作跟随与轨迹精度,支持受限内存下的分钟级连贯流式推演,且策略内对齐持续降低动作漂移与长时程不一致。

原文摘要 · Abstract (English)

Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and robot learning. Beyond clip-level future prediction, a unified generative formulation should relate actions to consequences, support flexible horizons and continuous interaction, and enable reward-driven optimization. We introduce WALL-SS, a world model that generates visual futures through Scale-wise autoregressive Scaling, enabling action-controllable and long-horizon robotic simulation. WALL-SS represents embodied trajectories as causal sequences of temporally interleaved observations and actions, making action-dependent state transitions explicit while naturally supporting variable-length generation, streaming extension through reusable causal states, and direct optimization through sequence probabilities. To make this formulation effective over long horizons, we generate each future observation in a coarse-to-fine manner and develop three complementary components within the same hierarchy. Action-conditioned next-scale prediction injects scale-aligned action representations to improve action-future coupling and model both successful and failed behaviors. Scale-compressed long-horizon memory retains recent interactions at fine resolution while compressing distant observations and actions, with scale-wise dream forcing enhancing robustness to self-generated context. Finally, on-policy alignment optimizes autoregressive visual dynamics with action-following and long-term consistency rewards while preserving the pretrained visual distribution. Experiments show that WALL-SS improves action following and trajectory accuracy, supports coherent minute-long streaming rollout under bounded memory, and consistently benefits from on-policy alignment in reducing action drift and long-horizon inconsistency.

世界模型机器人学习长时序生成自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。