用潜在运动链统一世界模型与动作推理,提升机器人视觉-语言-动作学习效率。
Chain of World: World Model Thinking in Latent Motion
- 通过预训练视频VAE分离视频结构与运动潜变量,构建连续潜在运动链。
- 在仿真环境中优于现有世界模型和潜在动作方法,实现高效视觉运动学习。
- 适合研究具身智能、视觉-语言-动作模型及机器人控制的开发者参考。
视觉-语言-动作(VLA)模型是实现具身智能的有前途路径,但常忽略视觉动态背后的预测性与时序因果结构。世界模型型VLA通过预测未来帧来弥补此缺陷,却浪费资源重建冗余背景。潜在动作型VLA紧凑编码帧间变化,但缺乏连续动态建模与世界知识。为此,本文提出CoWVLA(Chain-of-World VLA),一种新范式,将世界模型时序推理与解耦潜在运动表示相结合。首先,使用预训练视频变分自编码器(video VAE)作为潜在运动提取器,显式将视频片段分解为结构与运动潜变量。其次,在预训练阶段,VLA从指令与初始帧推断连续潜在运动链,并预测片段终点帧。最后,在联合微调阶段,通过统一自回归解码器联合建模稀疏关键帧与动作序列,对齐潜在动态与离散动作预测。该设计保留了世界模型的时序推理与世界知识优势,同时保持潜在动作的紧凑性与可解释性,支持高效视觉运动学习。大量实验在机器人仿真基准上表明,CoWVLA优于现有世界模型与潜在动作方法,具备中等计算效率,展现出更优的VLA预训练潜力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. Latent-action VLAs encode frame-to-frame transitions compactly, but lack temporally continuous dynamic modeling and world knowledge. To overcome these limitations, we introduce CoWVLA (Chain-of-World VLA), a new "Chain of World" paradigm that unifies world-model temporal reasoning with a disentangled latent motion representation. First, a pretrained video VAE serves as a latent motion extractor, explicitly factorizing video segments into structure and motion latents. Then, during pre-training, the VLA learns from an instruction and an initial frame to infer a continuous latent motion chain and predict the segment's terminal frame. Finally, during co-fine-tuning, this latent dynamic is aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder. This design preserves the world-model benefits of temporal reasoning and world knowledge while retaining the compactness and interpretability of latent actions, enabling efficient visuomotor learning. Extensive experiments on robotic simulation benchmarks show that CoWVLA outperforms existing world-model and latent-action approaches and achieves moderate computational efficiency, highlighting its potential as a more effective VLA pretraining paradigm. The project website can be found at https://fx-hit.github.io/cowvla-io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。