arXiv:2605.30957cs.RO2026-05

用强化学习生成高质量机器人演示,替代人力操作,提升训练效果。

RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning

论文配图:RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning
图 1 · 摘自论文原文
  • 用强化学习在仿真中生成结构化轨迹,再迁移到真实机器人
  • 生成轨迹更平滑,下游视觉语言模型性能优于人工示范
  • 适合需要大量高质量数据的机器人学习研究者

视觉-语言-动作(VLA)模型已成为通用机器人控制的有前景范式,但其性能受限于高质量机器人轨迹数据的获取。当前主要依赖人力遥控收集数据,成本高且难以扩展。本文提出RDGen,一种基于强化学习的仿真到现实演示生成框架。该框架将训练好的强化学习策略作为结构化轨迹生成器,结合基于VLM的任务解析模块、Grounding DINO物体定位模块,以及从仿真迁移到真实机器人的强化学习策略。成功轨迹被采集为高质量演示用于下游VLA训练,仿真阶段还能以极低边际成本提供大量额外轨迹。在抓取放置任务上的实验表明,迁移后的强化学习策略取得高任务成功率。与人工遥控相比,RDGen生成的轨迹更平滑,下游VLA模型表现更优。结果表明,强化学习生成的演示可作为更可靠、一致的监督信号,用于机器人策略学习。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robot control. However, their performance remains fundamentally constrained by the availability of high-quality robot trajectory data. In current robot learning practice, such data are primarily collected through human teleoperation, which is labor-intensive, costly, and difficult to scale. In this paper, we propose RDGen, a sim-to-real reinforcement learning framework for generating high-quality robot demonstrations. Rather than employing reinforcement learning solely as the final control policy, RDGen leverages trained RL policies as a structured trajectory generator. The system consists of a VLM-based task parser that identifies task-relevant objects, a Grounding DINO-based object localizer, and an RL policy transferred from simulation to the real robot. Successful rollouts are then harvested as clean, high-quality demonstrations for downstream VLA training, while the simulation stage further provides a scalable source of additional trajectories at little marginal cost. Experiments on a pick-and-place task demonstrate that the transferred RL policy achieves a high task success rate. Compared with human teleoperation, RDGen produces significantly smoother trajectories and yields superior downstream VLA performance. These results indicate that RL-generated demonstrations can serve as more reliable and consistent supervisory signals for robot policy learning.

机器人学习强化学习轨迹生成仿真到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。