用世界动作先验提升视觉语言动作模型的泛化能力。
World Pilot: Steering Vision-Language-Action Models with World-Action Priors

- 引入世界-动作先验,通过潜空间与动作双路径增强决策
- 在LIBERO-Plus零样本跨域任务上达到84.7%成功率,领先现有方法
- 适用于视角、几何、形变等多类变化场景,尤其适合真实机器人部署
视觉-语言-动作(VLA)模型依赖大规模预训练获得语义理解,在分布内操作任务中表现良好。然而,这种理解基于静态图像-文本对,无法捕捉操作过程中连续且接触丰富的动态特性。本文提出World Pilot框架,通过世界-动作模型(WAM)为策略引入先验信息,并经由两条互补路径注入决策链:潜空间引导使感知层基于场景演化潜变量进行条件化,动作引导则向动作生成器提供预期轨迹作为运动先验。两者共同赋予VLA对场景的预测视图和轨迹级运动提示,强化其在语义条件之外的前瞻性能力。即使使用未经过动作后训练的视频预训练世界模型,场景演化先验仍保持有效。World Pilot在LIBERO-Plus零样本跨域基准上实现84.7%的总成功率,是当前最佳表现,并在四个操作任务的真实机器人设置中均取得最高成功率,尤其在视角、几何、可变形状态和姿态变化下优势显著。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models inherit semantic grounding from large-scale pretraining and perform competently across in-distribution manipulation tasks. This grounding, however, is built on static image-text pairs, whereas manipulation is a continuous, contact-rich process whose dynamics this pretraining cannot capture. We present World Pilot, a VLA framework that augments the policy with priors from a World-Action Model (WAM), routed into the decision chain through two complementary pathways. Latent Steering conditions the perception layer on a scene-evolution latent, and Action Steering supplies an anticipated trajectory as a motion prior to the action generator. Together the two priors equip the VLA with an anticipated view of the scene and a trajectory-level motion hint alongside its semantic conditioning, and the scene-evolution prior remains effective even when supplied by a video-pretrained world model that has not been action-post-trained. World Pilot attains a state-of-the-art Total success rate of 84.7% on the LIBERO-Plus zero-shot OOD benchmark and the highest success rate on every real-robot setting across four manipulation tasks, with the largest margins under shifts in viewpoint, geometry, deformable state, and pose. Project Website: https://world-pilot.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。