arXiv:2605.30226cs.ROcs.AI2026-05

BORA让机器人在真实世界中更可靠地完成精细操作。

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

论文配图:BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models
图 1 · 摘自论文原文
  • 用视觉语言模型认知+动作片段联合评估动作好坏,实现离线价值引导。
  • 在线阶段引入人工干预的残差修正,成功率提升33%,泛化能力提高43%。
  • 适合需要高精度操控的机器人应用,尤其擅长处理真实环境差异。

视觉-语言-动作(VLA)模型为将视觉语言理解转化为现实机器人操作提供了新范式。然而,由于手部控制维度高且执行误差累积,精细操作仍具挑战性,真实世界强化学习后训练成为弥合视觉动作生成与物理可靠执行差距的关键。但高维精细探索常导致时间不一致、样本效率低及硬件风险。为此,我们提出BORA:一种面向真实世界精细操作的离线到在线强化学习后训练框架。离线阶段,BORA构建一个以视觉语言模型认知标记和动作片段为输入的评论家,实现动作条件化的价值引导,使评论家能超越视觉上下文评估精细手部动作。在线阶段,冻结VLA主干,引入轻量级人机协同(HiL)分段残差适应机制,纠正实际物理环境中的执行误差并优化离线学习的意图。通过继承离线评论家并采用干预驱动奖励,BORA有效修正执行偏差,适应真实物理变异,同时保留预训练策略作为稳定先验。在五个复杂真实任务上的大量实验表明,BORA显著优于纯模仿学习和传统解耦强化学习基线,在标准设置下平均成功率提升33%,未见物体泛化性能最高提升43%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains challenging for VLA policies due to high-dimensional hand control and compounding execution errors, which makes real-world RL post-training essential for bridging the gap between visually grounded action generation and physically reliable dexterous execution. However, high-dimensional dexterous exploration often triggers temporal inconsistency, sample inefficiency and hardware risks in the real world. To address these challenges, we propose BORA, an offline-to-online RL post-training framework designed for real-world dexterous VLA models. In the offline phase, BORA constructs a critic that takes both the VLM's cognition tokens and action chunks as inputs. This design enables action-conditioned value guidance, allowing the critic to evaluate dexterous hand motions beyond visual context alone. During the subsequent online phase, BORA freezes the VLA base and introduces a lightweight, Human-in-the-Loop (HiL) chunk-wise residual adaptation mechanism to mitigate real-world execution errors and further correct the offline-learned intents within the actual physical environment. By inheriting the offline critic and employing intervention-driven rewards, BORA effectively corrects execution discrepancies and adapts to real-world physical variances while preserving the pretrained policy as a stable prior. Extensive evaluations across five complex real-world dexterous tasks demonstrate that BORA significantly outperforms pure imitation learning and traditional decoupled RL baselines, achieving a 33% absolute increase in average success rate under standard settings and up to a 43% improvement in unseen object generalization.

机器人强化学习精细操作视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。