用模拟数据生成高质量机器人策略,无需人工标注
ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors
- 用扩散模型初始化不完美演示,再通过强化学习优化初始噪声
- 工业装配任务成功率90.5%,长程操作达85%,优于所有基线
- 适合需要高可靠性的真实机器人部署,尤其擅长复杂操作
学习可泛化且鲁棒的行为克隆策略需要大量高质量机器人数据。尽管人类示范(如遥操作)是专家行为的标准来源,但在现实世界中大规模获取此类数据代价高昂。本文提出ExpertGen框架,通过模拟自动化专家策略学习,实现可扩展的模拟到现实迁移。ExpertGen首先利用在不完美示范上训练的扩散模型初始化行为先验,这些示范可由大语言模型生成或人工提供。随后,通过强化学习优化扩散模型的初始噪声,在保持原始策略冻结的前提下,引导其向高任务成功率演进。由于预训练扩散模型保持冻结,ExpertGen将探索限制在安全、类人行为流形内,同时仅需稀疏奖励即可有效学习。在挑战性操纵基准测试中,ExpertGen稳定生成高质量专家策略,无需奖励工程。在工业装配任务中,整体成功率达到90.5%;在长时程操纵任务中达到85%,优于所有基线方法。所得策略表现出灵巧控制能力,并在多种初始配置和故障状态下保持鲁棒性。为验证模拟到现实的迁移效果,所学的状态基专家策略通过DAgger进一步蒸馏为视觉-运动策略,并成功部署于真实机器人硬件上。
原文摘要 · Abstract (English)
Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (e.g., through teleoperation) serve as the standard source for expert behaviors, acquiring such data at scale in the real world is prohibitively expensive. This paper introduces ExpertGen, a framework that automates expert policy learning in simulation to enable scalable sim-to-real transfer. ExpertGen first initializes a behavior prior using a diffusion policy trained on imperfect demonstrations, which may be synthesized by large language models or provided by humans. Reinforcement learning is then used to steer this prior toward high task success by optimizing the diffusion model's initial noise while keep original policy frozen. By keeping the pretrained diffusion policy frozen, ExpertGen regularizes exploration to remain within safe, human-like behavior manifolds, while also enabling effective learning with only sparse rewards. Empirical evaluations on challenging manipulation benchmarks demonstrate that ExpertGen reliably produces high-quality expert policies with no reward engineering. On industrial assembly tasks, ExpertGen achieves a 90.5% overall success rate, while on long-horizon manipulation tasks it attains 85% overall success, outperforming all baseline methods. The resulting policies exhibit dexterous control and remain robust across diverse initial configurations and failure states. To validate sim-to-real transfer, the learned state-based expert policies are further distilled into visuomotor policies via DAgger and successfully deployed on real robotic hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。