用强化学习生成高质量机器人演示,替代人力操作,提升训练效果。
RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning

- 用强化学习在仿真中生成结构化轨迹,再迁移到真实机器人
- 生成轨迹更平滑,下游视觉语言模型性能优于人工示范
- 适合需要大量高质量数据的机器人学习研究者
视觉-语言-动作(VLA)模型已成为通用机器人控制的有前景范式,但其性能受限于高质量机器人轨迹数据的获取。当前主要依赖人力遥控收集数据,成本高且难以扩展。本文提出RDGen,一种基于强化学习的仿真到现实演示生成框架。该框架将训练好的强化学习策略作为结构化轨迹生成器,结合基于VLM的任务解析模块、Grounding DINO物体定位模块,以及从仿真迁移到真实机器人的强化学习策略。成功轨迹被采集为高质量演示用于下游VLA训练,仿真阶段还能以极低边际成本提供大量额外轨迹。在抓取放置任务上的实验表明,迁移后的强化学习策略取得高任务成功率。与人工遥控相比,RDGen生成的轨迹更平滑,下游VLA模型表现更优。结果表明,强化学习生成的演示可作为更可靠、一致的监督信号,用于机器人策略学习。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robot control. However, their performance remains fundamentally constrained by the availability of high-quality robot trajectory data. In current robot learning practice, such data are primarily collected through human teleoperation, which is labor-intensive, costly, and difficult to scale. In this paper, we propose RDGen, a sim-to-real reinforcement learning framework for generating high-quality robot demonstrations. Rather than employing reinforcement learning solely as the final control policy, RDGen leverages trained RL policies as a structured trajectory generator. The system consists of a VLM-based task parser that identifies task-relevant objects, a Grounding DINO-based object localizer, and an RL policy transferred from simulation to the real robot. Successful rollouts are then harvested as clean, high-quality demonstrations for downstream VLA training, while the simulation stage further provides a scalable source of additional trajectories at little marginal cost. Experiments on a pick-and-place task demonstrate that the transferred RL policy achieves a high task success rate. Compared with human teleoperation, RDGen produces significantly smoother trajectories and yields superior downstream VLA performance. These results indicate that RL-generated demonstrations can serve as more reliable and consistent supervisory signals for robot policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。