arXiv:2512.16861cs.ROcs.AI2025-12被引 4

用模仿学习+强化学习,让机器人自动完成复杂长时序操作任务

ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning

  • 将任务分解为局部技能,通过运动规划连接并训练
  • 在Robosuite上达80%成功率,优化后性能提升89%
  • 适合需要长序列动作的机器人控制研究者

长时序操作是机器人领域的长期挑战。我们提出ReinforceGen,结合任务分解、数据生成、模仿学习与运动规划构建初始方案,并通过基于强化学习的在线微调优化各组件。系统首先将任务拆分为多个局部技能,通过运动规划连接;技能与规划目标在10个真人示范生成的数据集上进行模仿学习训练,随后通过在线适应与强化学习进一步优化。在Robosuite数据集上,使用视觉-运动控制在最高重置范围设置下达到80%成功率。消融实验表明,微调方法使平均性能提升89%。真实场景评估也验证了微调带来的显著改进。更多结果与视频见https://reinforcegen.github.io。

原文摘要 · Abstract (English)

Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines task decomposition, data generation, imitation learning, and motion planning to form an initial solution, and improves each component through reinforcement-learning-based fine-tuning. ReinforceGen first segments the task into multiple localized skills, which are connected through motion planning. The skills and motion planning targets are trained with imitation learning on a dataset generated from 10 human demonstrations, and then fine-tuned through online adaptation and reinforcement learning. When benchmarked on the Robosuite dataset, ReinforceGen reaches 80% success rate on all tasks with visuomotor controls in the highest reset range setting. Additional ablation studies show that our fine-tuning approaches contribute to an 89% average performance increase. Finally, ReinforceGen demonstrates significant improvement through fine-tuning in our real-world evaluations. More results and videos are available at https://reinforcegen.github.io.

机器人强化学习技能分解模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。