用视频帧间时间距离学习密集奖励,让机器人更快学会复杂任务。
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
- 通过分析视频帧间时间距离,自动提取任务进展信号
- 在10个任务中9个达到近100%成功率,仅需20万次交互
- 可利用人类视频预训练,适合缺乏标注数据的场景
设计密集奖励对强化学习至关重要,但在机器人领域常需大量人工工作且难以扩展。一种有前景的方法是将任务进展视为密集奖励信号,量化动作随时间推进系统完成任务的程度。我们提出TimeRewarder,一种简单而有效的奖励学习方法,通过建模帧对之间的时序距离,从被动视频(包括机器人示范和人类视频)中提取进度估计信号。实验表明,TimeRewarder可为强化学习提供逐步代理奖励。在十个具有挑战性的Meta-World任务上,其显著提升了稀疏奖励任务的强化学习性能,在每个任务仅20万次环境交互下,9/10任务达到接近完美的成功率。该方法优于以往方法,甚至超过人工设计的环境密集奖励,无论在最终成功率还是样本效率上。此外,我们还展示了TimeRewarder可通过真实世界人类视频进行预训练,凸显其从多样化视频源获取丰富奖励信号的可扩展潜力。
原文摘要 · Abstract (English)
Designing dense rewards is crucial for reinforcement learning (RL), yet in robotics it often demands extensive manual effort and lacks scalability. One promising solution is to view task progress as a dense reward signal, as it quantifies the degree to which actions advance the system toward task completion over time. We present TimeRewarder, a simple yet effective reward learning method that derives progress estimation signals from passive videos, including robot demonstrations and human videos, by modeling temporal distances between frame pairs. We then demonstrate how TimeRewarder can supply step-wise proxy rewards to guide reinforcement learning. In our comprehensive experiments on ten challenging Meta-World tasks, we show that TimeRewarder dramatically improves RL for sparse-reward tasks, achieving nearly perfect success in 9/10 tasks with only 200,000 environment interactions per task. This approach outperformed previous methods and even the manually designed environment dense reward on both the final success rate and sample efficiency. Moreover, we show that TimeRewarder pretraining can exploit real-world human videos, highlighting its potential as a scalable approach to rich reward signals from diverse video sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。