arXiv:2609.06251cs.CVcs.RO2026-09

让机器人更懂指令,实现精准长时任务执行

MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

论文配图:MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control
图 1 · 摘自论文原文
  • 通过强化学习显式耦合推理与控制,提升指令到动作的一致性
  • 在真实机器人上实现10.0分的抓取任务成功率提升,平均导航准确率高1.6点
  • 适用于移动、四足和人形机器人,支持跨平台通用控制

将自然语言指令转化为可靠可执行动作,仍是移动机器人视觉-语言-动作系统的核心挑战。现有方法依赖隐式推理或整体动作预测,难以兼顾长时决策连贯性与动作精度。为此,我们提出MobileVLA-R1 2.0,一种基于强化学习的增强型框架,显式连接结构化具身推理与可执行控制。该框架通过监督式思维链对齐与强化学习,在多粒度轨迹上学习推理能力,超越纯行为监督。为支持移动与操作任务,引入条件推理的动作解码器,将多模态推理表示映射为任务级动作目标,并由控制器转为具体执行命令。该设计构建统一感知-推理-动作接口,同时解耦高层动作生成与机器人特异性执行。我们在语言引导导航、四足控制与人形移动操作任务中评估,涵盖VLN-CE、QUARD及实际部署于Unitree Go2与G1机器人。MobileVLA-R1 2.0持续优于强基线模型,在VLN-CE上平均成功率达1.6分提升,在真实世界G1机械臂任务中全任务成功率提升10.0分,展现出鲁棒的长时指令遵循与闭环执行能力。

原文摘要 · Abstract (English)

Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semantic reasoning and low-level locomotion and manipulation control. Existing approaches often rely on implicit reasoning or monolithic action prediction, making it difficult to maintain coherent long-horizon decision making while producing precise and adaptable robot actions. To address this challenge, we propose MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control. The framework learns multi-granularity reasoning over embodied trajectories through supervised Chain-of-Thought (CoT) alignment and reinforcement learning, improving reasoning-to-action consistency beyond purely behavioral supervision. To support both locomotion and manipulation, we further introduce a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers. This design provides a unified perception-reasoning-action interface while decoupling high-level action generation from robot-specific actuation. We conduct extensive evaluations on language-guided navigation, quadruped control, and humanoid mobile manipulation, covering VLN-CE, QUARD, and real-world deployments on Unitree Go2 and G1 robots. MobileVLA-R1 2.0 consistently outperforms strong VLA baselines, achieving an average 1.6 point improvement in SR on VLN-CE and a 10.0 point improvement in full-task success on real-world G1 mobile manipulation tasks over MobileVLA-R1, while demonstrating robust long-horizon instruction following and closed-loop execution across different robotic platforms.

机器人控制语言理解强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。