用文本到视频扩散模型生成密集奖励,提升机器人操控的强化学习效率
TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning
- 用预训练文本到视频扩散模型预测动作序列,对比当前观测生成密集奖励
- 在11个复杂任务中优于传统稀疏奖励方法,样本效率和性能更优
- 适合需要高效训练的机器人操控场景,无需真实环境奖励
为构建通用智能体,特别是在机器人操作这一挑战性领域,发展可扩展且泛化的强化学习(RL)奖励工程至关重要。尽管近期基于视觉-语言模型(VLMs)的奖励设计已显成效,但其稀疏奖励特性严重限制了样本效率。本文提出TeViR,一种新方法:利用预训练的文本到视频扩散模型生成密集奖励,通过比较模型预测的图像序列与当前观测来实现。在11个复杂机器人任务上的实验表明,TeViR优于依赖稀疏奖励的传统方法及其他最先进(SOTA)方法,在无真实环境奖励条件下仍实现更高的样本效率和性能。TeViR在复杂环境中高效引导智能体的能力,凸显其在机器人操控强化学习应用中的巨大潜力。
原文摘要 · Abstract (English)
Developing scalable and generalizable reward engineering for reinforcement learning (RL) is crucial for creating general-purpose agents, especially in the challenging domain of robotic manipulation. While recent advances in reward engineering with Vision-Language Models (VLMs) have shown promise, their sparse reward nature significantly limits sample efficiency. This paper introduces TeViR, a novel method that leverages a pre-trained text-to-video diffusion model to generate dense rewards by comparing the predicted image sequence with current observations. Experimental results across 11 complex robotic tasks demonstrate that TeViR outperforms traditional methods leveraging sparse rewards and other state-of-the-art (SOTA) methods, achieving better sample efficiency and performance without ground truth environmental rewards. TeViR's ability to efficiently guide agents in complex environments highlights its potential to advance reinforcement learning applications in robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。