让机器人学会不用视觉也能完成复杂操作,且适应新任务。
VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation
- 用真实世界强化学习,教师教学生从有视觉到无视觉的技能迁移。
- 50分钟训练后在3个任务上成功率95%,8个未见任务也表现稳定。
- 适合追求真实场景鲁棒性与泛化能力的研究者或工程师。
在接触密集型机器人操作中,视觉可提供任务相关线索,加速学习进程。然而,依赖视觉的策略容易对训练时的视觉条件过拟合,降低鲁棒性和可迁移性。我们提出一种人机协同的强化学习框架,通过教师-学生知识蒸馏,在真实世界中完全训练,无需领域随机化或数据增强。视觉增强的教师将知识蒸馏给仅依赖位姿、速度和力矩感知的无视觉学生,实现快速训练与强泛化能力。在真实世界的NIST装配基准板上,该方法在约50分钟内完成3个代表性任务的训练,整体成功率达95%,并成功泛化至8个未见过的任务变体。通过蒸馏微调,最困难任务达到100%成功率。结果表明,所获策略在鲁棒性和适应性上均优于基线。
原文摘要 · Abstract (English)
When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning beyond what proprioception alone can achieve. However, vision-enabled policies tend to overfit to the visual conditions seen during training, limiting their robustness and transferability. We present a human-in-the-loop RL framework that employs teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation. A vision-enabled teacher distills its knowledge into a vision-free student that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization. On the real-world NIST assembly benchmark board, our approach achieves 95\% overall success after approximately 50 minutes of training on 3 representative tasks, including robust generalization to 8 unseen task variants. Fine-tuning with distillation achieves full success on the most challenging task. We demonstrate that the resulting policies outperform baselines in both robustness and adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。