arXiv:2603.12553cs.ROcs.CV2026-03

用结构化帧替代密集预测,让机器人规划更精准可靠。

Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation

  • 用物理有意义的结构化帧替代密集视觉预测,提升规划可解释性。
  • 在Sim-Env和LIBERO上分别达75.0%和94.8%成功率,长程任务表现优异。
  • 适合需要高可靠性与可泛化的复杂机械臂操控场景。

基于世界模型的视觉-语言-动作(VLA)架构通过预测视觉前景提升了机器人操作能力。然而,密集未来预测带来视觉冗余并累积误差,导致长程计划漂移。而现有稀疏方法多以高层语义子任务或隐式潜在状态表示前景,缺乏显式的运动学对齐,削弱了规划与底层执行的一致性。为此,我们提出StructVLA,将生成式世界模型重构为显式的结构化规划器。不同于密集滚动或语义目标,StructVLA预测稀疏但具物理意义的结构化帧,其由内在运动学线索(如夹爪状态切换、运动转折点)导出,紧密对齐任务进展。通过两阶段训练范式与统一离散标记词汇表实现:先训练模型预测结构化帧,再优化其映射至低层动作。该方法提供清晰物理引导,连接视觉规划与运动控制。实验表明,StructVLA在SimplerEnv-WidowX上平均成功率75.0%,在LIBERO上达94.8%。真实世界部署进一步验证了其在基础抓取与复杂长程任务中的可靠完成能力与强泛化性。

原文摘要 · Abstract (English)

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates errors, causing long-horizon plan drift. Meanwhile, recent sparse methods typically represent visual foresight using high-level semantic subtasks or implicit latent states. These representations often lack explicit kinematic grounding, weakening the alignment between planning and low-level execution. To address this, we propose StructVLA, which reformulates a generative world model into an explicit structured planner for reliable control. Instead of dense rollouts or semantic goals, StructVLA predicts sparse, physically meaningful structured frames. Derived from intrinsic kinematic cues (e.g., gripper transitions and kinematic turning points), these frames capture spatiotemporal milestones closely aligned with task progress. We implement this approach through a two-stage training paradigm with a unified discrete token vocabulary: the world model is first trained to predict structured frames and subsequently optimized to map the structured foresight into low-level actions. This approach provides clear physical guidance and bridges visual planning and motion control. In our experiments, StructVLA achieves strong average success rates of 75.0% on SimplerEnv-WidowX and 94.8% on LIBERO. Real-world deployments further demonstrate reliable task completion and robust generalization across both basic pick-and-place and complex long-horizon tasks.

机器人控制世界模型结构化规划视觉预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。