用视觉语言动作模型提升机器人现实世界强化学习效率
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
- 基于多模态数据训练通用奖励模型,自动评估任务进展
- 200次真实交互后成功率从30%提升至90%,人机协同再增50%效率
- 无需任务定制奖励函数,支持新任务快速适配
基于视觉-语言-动作(VLA)模型的机器人现实世界强化学习受限于稀疏的手工奖励和低效探索。本文提出VLAC,一个基于InternVL、在大规模异构数据上训练的通用过程奖励模型。给定成对观测与语言目标,它输出稠密的进展增量和完成信号,消除任务特定的奖励工程,并支持对未见任务和环境的一次性上下文迁移。VLAC在视觉-语言数据上训练以增强感知、对话与推理能力,结合机器人与人类轨迹数据以实现动作生成与进展估计的具身化,同时通过构建大量负样本与语义不匹配样本,强化其拒绝无关提示及检测退化或停滞的能力。通过提示控制,单一VLAC模型交替生成奖励与动作令牌,统一了评判者与策略。部署于异步真实世界强化学习循环中,结合分层人机协同协议(离线示范回放、返回探索、人工引导探索),显著加速探索并稳定早期学习。在四项不同现实操作任务中,VLAC将成功率从约30%提升至约90%,引入人机协同干预后样本效率进一步提升50%,最终成功率达100%。
原文摘要 · Abstract (English)
Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLAC, a general process reward model built upon InternVL and trained on large scale heterogeneous datasets. Given pairwise observations and a language goal, it outputs dense progress delta and done signal, eliminating task-specific reward engineering, and supports one-shot in-context transfer to unseen tasks and environments. VLAC is trained on vision-language datasets to strengthen perception, dialogic and reasoning capabilities, together with robot and human trajectories data that ground action generation and progress estimation, and additionally strengthened to reject irrelevant prompts as well as detect regression or stagnation by constructing large numbers of negative and semantically mismatched samples. With prompt control, a single VLAC model alternately generating reward and action tokens, unifying critic and policy. Deployed inside an asynchronous real-world RL loop, we layer a graded human-in-the-loop protocol (offline demonstration replay, return and explore, human guided explore) that accelerates exploration and stabilizes early learning. Across four distinct real-world manipulation tasks, VLAC lifts success rates from about 30\% to about 90\% within 200 real-world interaction episodes; incorporating human-in-the-loop interventions yields a further 50% improvement in sample efficiency and achieves up to 100% final success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。