arXiv:2604.23073cs.LGcs.RO2026-04被引 34

用轻量方法让视觉语言动作模型快速在线强化学习,提升机器人操作精度与速度。

RL Token: Bootstrapping Online RL with Vision-Language-Action Models

论文配图:RL Token: Bootstrapping Online RL with Vision-Language-Action Models
图 1 · 摘自论文原文
  • 在预训练模型中引入可优化的RL令牌,作为强化学习接口。
  • 仅需数小时真实世界训练,任务成功率显著提升,最快提速3倍。
  • 适合需要快速适配新任务的机器人系统,尤其对大模型高效微调有优势。

视觉-语言-动作(VLA)模型可直接学习多种操作技能,但要达到真实场景所需的精度和速度,仍需进一步微调,例如通过强化学习(RL)。我们提出一种轻量级方法,仅需数小时真实世界实践,即可实现预训练VLA的高效在线强化学习微调。首先,将VLA改造为暴露一个“RL令牌”——一种紧凑的读出表示,既能保留任务相关的预训练知识,又能作为在线强化学习的有效接口;其次,在该RL令牌上训练小型演员-评论家头,以优化动作,同时保持学习策略与VLA的一致性。使用RL令牌的在线强化学习(RLT)使大型VLA也能快速高效地进行强化学习微调。在四个真实机器人任务(螺钉安装、扎带绑扎、充电器插入、网线插入)中,RLT将最困难部分的执行速度最高提升3倍,并在几分钟至数小时内显著提高成功率。某些任务甚至超过人类遥操作的速度。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models can learn to perform diverse manipulation skills "out of the box," but achieving the precision and speed that real-world tasks demand requires further fine-tuning -- for example, via reinforcement learning (RL). We introduce a lightweight method that enables sample-efficient online RL fine-tuning of pretrained VLAs using just a few hours of real-world practice. We (1) adapt the VLA to expose an "RL token," a compact readout representation that preserves task-relevant pretrained knowledge while serving as an efficient interface for online RL, and (2) train a small actor-critic head on this RL token to refine the actions, while anchoring the learned policy to the VLA. Online RL with the RL token (RLT) makes it possible to fine-tune even large VLAs with RL quickly and efficiently. Across four real-robot tasks (screw installation, zip tie fastening, charger insertion, and Ethernet insertion), RLT improves the speed on the hardest part of the task by up to 3x and raises success rates significantly within minutes to a few hours of practice. It can even surpass the speed of human teleoperation on some of the tasks.

强化学习机器人VLA在线微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。