用在线拒收采样让机器人学得更稳更快,还能自动纠错。
Human-in-the-loop Online Rejection Sampling for Robotic Manipulation
- 训练时过滤负奖励样本,稳定价值估计
- 1.5小时实机训练即掌握复杂操作,性能超越基准方法
- 支持人机协同修正,适合需要高鲁棒性的实际场景
强化学习(RL)常用于生成稳健的机器人操作策略,但通过强化学习微调视觉-语言-动作(VLA)模型时,由于价值估计不准和中间步骤监督稀疏,易出现不稳定。相反,模仿学习(IL)虽易训练,却因离线特性表现较差。本文提出Hi-ORS,一种简单有效的后训练方法,利用拒绝采样实现训练稳定与高鲁棒性。该方法在在线微调中过滤负奖励样本以稳定价值估计,并采用奖励加权的监督训练目标,提供密集的中间步骤监督。为系统研究,我们构建了异步推理-训练框架,支持灵活的人机在线协同修正,作为学习错误恢复行为的显式指导。在三个真实任务和两种机器人平台上,Hi-ORS仅用1.5小时真实世界训练即可使pi-base策略掌握接触丰富的操作,其效果和效率显著优于RL与IL基线。值得注意的是,微调后的策略在测试时表现出强可扩展性,能可靠执行复杂错误恢复行为,从而获得更优表现。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is widely used to produce robust robotic manipulation policies, but fine-tuning vision-language-action (VLA) models with RL can be unstable due to inaccurate value estimates and sparse supervision at intermediate steps. In contrast, imitation learning (IL) is easy to train but often underperforms due to its offline nature. In this paper, we propose Hi-ORS, a simple yet effective post-training method that utilizes rejection sampling to achieve both training stability and high robustness. Hi-ORS stabilizes value estimation by filtering out negatively rewarded samples during online fine-tuning, and adopts a reward-weighted supervised training objective to provide dense intermediate-step supervision. For systematic study, we develop an asynchronous inference-training framework that supports flexible online human-in-the-loop corrections, which serve as explicit guidance for learning error-recovery behaviors. Across three real-world tasks and two embodiments, Hi-ORS fine-tunes a pi-base policy to master contact-rich manipulation in just 1.5 hours of real-world training, outperforming RL and IL baselines by a substantial margin in both effectiveness and efficiency. Notably, the fine-tuned policy exhibits strong test-time scalability by reliably executing complex error-recovery behaviors to achieve better performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。