让机器人模型学会预测动作未来动态,提升操作成功率。
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- 用双头结构联合训练动作与运动图像扩散,增强预测能力。
- 在LIBERO上成功率达97.5%,真实场景性能提升23%。
- 无需改变推理路径,适合大规模视觉语言动作模型部署。
视觉-语言-动作(VLA)模型通过直接将多模态观测和指令映射为动作,在机器人操作中取得了显著进展。然而,它们通常仅模仿专家轨迹,缺乏对动作的预测性运动推理能力,限制了其决策能力。为此,本文提出联合学习运动图像扩散策略,通过在标准VLA架构中引入双头设计,其中动作头仍预测动作块,新增的运动头采用扩散变换器(DiT)预测基于光流的运动图像,以捕捉未来动态。两头联合训练,使共享的视觉语言模型(VLM)骨干网络学习到融合机器人控制与运动知识的表示。该方法在不改变标准VLA推理路径的前提下,构建了时序连贯且物理合理的表征,保持测试延迟不变。仿真与真实环境实验表明,该方法使pi-series VLA在LIBERO基准上成功率达到97.5%,在RoboTwin基准上达到58.0%,真实世界性能提升23%,验证了其有效提升大规模VLA运动推理能力的能力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they typically mimic expert trajectories without predictive motion reasoning, which limits their ability to reason about what actions to take. To address this limitation, we propose joint learning with motion image diffusion, a novel strategy that enhances VLA models with motion reasoning capabilities. Our method extends the VLA architecture with a dual-head design: while the action head predicts action chunks as in vanilla VLAs, an additional motion head, implemented as a Diffusion Transformer (DiT), predicts optical-flow-based motion images that capture future dynamics. The two heads are trained jointly, enabling the shared VLM backbone to learn representations that couple robot control with motion knowledge. This joint learning builds temporally coherent and physically grounded representations without modifying the inference pathway of standard VLAs, thereby maintaining test-time latency. Experiments in both simulation and real-world environments demonstrate that joint learning with motion image diffusion improves the success rate of pi-series VLAs to 97.5% on the LIBERO benchmark and 58.0% on the RoboTwin benchmark, yielding a 23% improvement in real-world performance and validating its effectiveness in enhancing the motion reasoning capability of large-scale VLAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。