用扩散强化学习生成高质量机器人操作数据,提升视觉语言动作模型性能。
Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
- 基于扩散模型优化策略,生成平滑且一致的动作轨迹。
- 在LIBERO基准上,自动生成数据训练的模型成功率81.9%,优于人类数据+5.3%。
- 适合需要大量高质量演示数据的机器人学习研究者使用。
视觉-语言-动作(VLA)模型在跨任务和机器人平台间表现出强泛化能力,但其依赖大规模人工示范,导致数据收集成本高昂。强化学习(RL)可自主生成示范,但传统算法在长时程、稀疏奖励的操作任务中表现不佳。本文提出一种改进的扩散策略优化算法,用于生成高质量、低方差的动作轨迹,构建基于扩散强化学习的VLA训练流程。该方法兼具扩散模型对复杂多样行为的表达能力,以及迭代去噪过程带来的隐式正则化,使生成轨迹更平滑一致。在包含130个长时程操作任务的LIBERO基准上评估,生成轨迹比人类示范和标准高斯RL策略更平滑、更一致。仅使用扩散RL生成数据训练的VLA模型平均成功率达81.9%,高于人类数据训练模型的+5.3%,也优于高斯RL生成数据训练模型的+12.6%。结果表明,扩散强化学习是生成海量高质量、低方差示范数据的有效替代方案。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have shown strong generalization across tasks and embodiments; however, their reliance on large-scale human demonstrations limits their scalability owing to the cost and effort of manual data collection. Reinforcement learning (RL) offers a potential alternative to generate demonstrations autonomously, yet conventional RL algorithms often struggle on long-horizon manipulation tasks with sparse rewards. In this paper, we propose a modified diffusion policy optimization algorithm to generate high-quality and low-variance trajectories, which contributes to a diffusion RL-powered VLA training pipeline. Our algorithm benefits from not only the high expressiveness of diffusion models to explore complex and diverse behaviors but also the implicit regularization of the iterative denoising process to yield smooth and consistent demonstrations. We evaluate our approach on the LIBERO benchmark, which includes 130 long-horizon manipulation tasks, and show that the generated trajectories are smoother and more consistent than both human demonstrations and those from standard Gaussian RL policies. Further, training a VLA model exclusively on the diffusion RL-generated data achieves an average success rate of 81.9%, which outperforms the model trained on human data by +5.3% and that on Gaussian RL-generated data by +12.6%. The results highlight our diffusion RL as an effective alternative for generating abundant, high-quality, and low-variance demonstrations for VLA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。