arXiv:2606.00985cs.RO2026-06

用运动规划增强VLA模型,无需更多数据就能提升长程任务鲁棒性

Make Your VLA More Robust Without More Data By Interleaving Motion Planning

论文配图:Make Your VLA More Robust Without More Data By Interleaving Motion Planning
图 1 · 摘自论文原文
  • 将视觉语言动作模型与基于模型的运动规划交替使用
  • 在BEHAVIOR-1K上任务进展提升113%优于顶尖端到端VLA
  • 适合做移动操作的机器人系统研发人员参考

视觉-语言-动作(VLA)模型在移动操作任务中表现优异,但在长程任务上的性能仍不理想。这主要源于两点:一是高阶目标需贯穿空间分散的多个子任务持续推进;二是早期执行错误会随任务时长快速累积。即使在大规模人类遥控数据上微调,问题依然存在,表明仅增加数据无法根本解决。为此,我们提出MPVI:运动规划与VLA的交替框架,通过融合基于模型的运动规划与VLA,实现无需额外训练即可提升鲁棒性。该方法支持在杂乱场景中通过开放词汇目标检测、前沿探索和运动规划定位并导航至远距离或被遮挡的目标物体。集成关键在于模块间可靠切换,我们通过基于视觉语言模型的完成状态检测与本体感知触发实现。在BEHAVIOR-1K基准测试中,该方法相较顶尖端到端VLA基线任务进展提升113%。更多信息见项目页:https://mpvi.netlify.app/

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown remarkable progress for mobile manipulation, but their performance on long-horizon tasks remains poor. These tasks are especially challenging because (1) progress toward high-level goals must be maintained across extended sequences of spatially distributed subtasks, and (2) early execution errors compound rapidly over the task horizon. These challenges persist despite finetuning on large human teleoperated mobile manipulation data, indicating that more data alone may not resolve the problem. To address these challenges, we propose MPVI: Motion Planner / VLA Interleaving, a framework that integrates model-based motion planning with VLAs to improve robustness without further training. The proposed integration enables localization and navigation to distant or occluded target objects through cluttered scenes using open-vocabulary object detection, frontier exploration and motion planning. However, such integration is non-trivial, requiring reliable switching between modules; we show one way forward via VLM-based completion checking with proprioceptive triggers. We evaluate our approach on the BEHAVIOR-1K benchmark and demonstrate 113% improvement in task progress over a top end-to-end VLA baseline. Additional details are available at the project page: https://mpvi.netlify.app/.

VLA运动规划机器人长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。