arXiv:2412.09858cs.ROcs.AI2024-12被引 51

用强化学习生成高质量数据,提升机器人通用策略的操控成功率。

RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning

  • 用强化学习生成任务特定数据,优化通用策略训练
  • 实测成功率最高提升40%,泛化能力更强
  • 适合需要高精度操作的机器人系统研发者

近期机器人基础模型的发展催生了可适应多种任务的通用策略。然而其性能高度依赖训练数据质量。本文提出强化学习蒸馏通用策略(RLDG),通过强化学习生成高质量训练数据以微调通用策略。在连接器插入、装配等精准操作任务的大量真实实验中,使用RL生成数据训练的策略显著优于人类示范数据训练的策略,成功率最高提升40%,且对新任务泛化能力更强。详细分析表明,性能提升源于更优的动作分布和更广的状态覆盖。结果表明,将任务特定强化学习与通用策略蒸馏结合,是兼具灵活性与高性能的机器人操作系统的有效路径。视频与代码见项目网站 https://generalist-distillation.github.io

原文摘要 · Abstract (English)

Recent advances in robotic foundation models have enabled the development of generalist policies that can adapt to diverse tasks. While these models show impressive flexibility, their performance heavily depends on the quality of their training data. In this work, we propose Reinforcement Learning Distilled Generalists (RLDG), a method that leverages reinforcement learning to generate high-quality training data for finetuning generalist policies. Through extensive real-world experiments on precise manipulation tasks like connector insertion and assembly, we demonstrate that generalist policies trained with RL-generated data consistently outperform those trained with human demonstrations, achieving up to 40% higher success rates while generalizing better to new tasks. We also provide a detailed analysis that reveals this performance gain stems from both optimized action distributions and improved state coverage. Our results suggest that combining task-specific RL with generalist policy distillation offers a promising approach for developing more capable and efficient robotic manipulation systems that maintain the flexibility of foundation models while achieving the performance of specialized controllers. Videos and code can be found on our project website https://generalist-distillation.github.io

机器人强化学习通用策略数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。