让视觉语言动作模型在真实世界中高效自适应,提升精准操作能力。
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

- 采用异步协同自举机制,分阶段优化策略与价值估计。
- 9个高精度化学任务平均成功率98.3%,单任务仅需45.8分钟。
- 适合需要高重复性、实时优化的机器人精准操控场景。
预训练的视觉-语言-动作(VLA)模型虽具备广泛操作能力,但在要求高精度与可重复性的任务中仍不可靠。将真实世界的在线强化学习(RL)应用于VLA后训练,可实现超越示范数据的自主试错优化,但面临两大瓶颈:1)不可靠的价值信号易引发策略漂移;2)大模型开销限制吞吐量与样本效率。为此,我们提出VLA-Precision框架,包含异步协同自举(ACoB)算法与ACoB-Stream架构。ACoB通过跨时尺度的不对称自举:早期干预引导的行为学习快速提升策略性能并改善在线经验质量;随着自主经验积累,全局回报传播与局部偏好排序逐步校准价值估计,生成相对动作优势用于参考正则化策略改进,抑制漂移。为在大型VLA上实现ACoB,我们设计了闭环经验-策略架构ACoB-Stream,以不变状态解耦和按需流式处理为核心原则,实现最高10.9×的吞吐量与计算效率提升。在四个机器人形态、九类高精度化学任务上的评估显示,VLA-Precision实现98.3%的平均成功率,每任务仅需45.8分钟,单轮轨迹运行速度达基线的1.2×至1.8×。资源见https://vla-precision.github.io。
原文摘要 · Abstract (English)
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral learning rapidly improves policy performance while enhancing online experience quality. As autonomous experience accumulates, global return propagation and local preference ranking progressively calibrate value estimates, yielding relative action advantages for reference-regularized policy improvement while suppressing drift. To enable ACoB on large VLAs, we develop ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, delivering up to 10.9$\times$ improvements in throughput and computational efficiency. Extensive evaluations on nine high-precision chemistry tasks across four categories and four robot embodiments show that VLA-Precision achieves 98.3\% mean success rate in 45.8 min/task, with 27.6 s episodes running at 1.2$\times$ and 1.8$\times$ the speeds of VLA and RL baselines. Resources are available at https://vla-precision.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。