arXiv:2608.01013cs.RO2026-08

用强化学习零样本适配新机器人,让语言指令控制全新机械臂。

RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment

  • 用仿真中密集几何奖励进行两阶段强化学习,先学方向动作再扩展物体指令
  • 指令成功率从34.25%提升至53.50%,对'左移''后退'效果显著
  • 无需特定数据集,适合无演示的新型机器人快速部署

将预训练视觉-语言-动作(VLA)策略适配到新机器人通常依赖特定形态的示范数据,这对形态迥异于主流机器人数据集的定制化机器人尤为受限。本文研究更难的零示范形态对齐问题:在具有简单夹爪和未见控制接口的缆索驱动并联机器人(CDPR)上,使用OpenVLA-OFT模型。不采用监督微调,而是通过仿真中的密集几何奖励进行强化学习。训练分两阶段:第一阶段用PPO学习方向动作;第二阶段从PPO检查点继续,采用扩展指令空间(包含物体条件指令)的GRPO。在四个共享方向指令上,平均保留测试成功率从PPO后的34.25%提升至PPO→GRPO后的53.50%,尤其在'move left'和'move backward'上提升明显。在GRPO阶段新增八种目标物体的'move to <object>'指令,获得39/400=9.75%严格成功,定性滚动显示多数情况能正确朝向目标,但后期存在不稳定现象。相比依赖示范数据集与标准刚性臂形态的先前OpenVLA及OpenVLA-OFT方法,本方法完全无需特定形态数据集。结果虽未建立鲁棒操作能力,但有力表明仅靠强化学习可为真正新颖的机器人形态生成首个可用的语言控制程序。

原文摘要 · Abstract (English)

Adapting a pretrained vision-language-action (VLA) policy to a new robot usually assumes embodiment-specific demonstrations. This assumption is especially restrictive for custom robots whose morphology differs strongly from the manipulators seen in large robot datasets. We study a harder setting: zero-demo embodiment alignment of OpenVLA-OFT on a cable-driven parallel robot (CDPR) with a simple gripper and a previously unseen control interface. Instead of supervised fine-tuning, we use reinforcement learning in simulation with dense geometric rewards computed from simulator state. The training is performed in two stages: a PPO stage for directional motion primitives, followed by GRPO continuation from the PPO checkpoint with an expanded instruction space that includes object-conditioned commands. On the four shared directional instructions, the average held-out success rate improves from 34.25\% after PPO to 53.50\% after PPO$\rightarrow$GRPO, with especially large gains on \texttt{move left} and \texttt{move backward}. In the GRPO stage we additionally introduce \texttt{move to <object>} over eight target objects and obtain 39/400 = 9.75\% strict success, while qualitative rollouts frequently show correct target-directed approach behavior before late-stage instability. Compared with prior OpenVLA and OpenVLA-OFT results, which rely on demonstration datasets and mostly standard rigid-arm embodiments, our method uses no embodiment-specific dataset at all. The results do not yet establish robust manipulation, but they provide stronger evidence that RL-only bootstrapping can create the first usable language-conditioned controller for a genuinely novel embodiment.

强化学习机器人控制零样本语言指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。