arXiv:2503.18210cs.LGcs.AI2025-03被引 2

用视频数据训练价值函数,引导在线强化学习高效探索。

ViVa: Video-Trained Value Functions for Guiding Online RL from Diverse Data

  • 从互联网视频等多样化数据中学习目标相关价值函数
  • 预训练后在新任务上实现跨目标泛化,且性能随数据量提升
  • 无需专家标注,适合缺乏奖励信号的复杂环境

在线强化学习在稀疏奖励场景下面临挑战,主要源于缺乏通往目标状态的有效反馈。由于带奖励信号的专家离线数据罕见,难以用于引导在线学习。如何在无任务特定数据情况下引导智能体?奖励塑形通过提供细粒度信号,可有效引导策略向最优解靠近。但传统方法依赖领域知识手动设计启发式规则。为此,我们提出一种数据驱动的方法,利用广泛可用的视频数据(如网络录像、非任务示范、失败演示和无目的交互)自动学习通用引导信号。通过构建从多样化被动数据中学习的意图条件价值函数,并将其融入奖励函数,我们实现了对在线强化学习的有效引导。实验表明,该方法适用于多种数据源,具有正向迁移能力(尤其在人类视频预训练后),能泛化至未见目标,且性能随数据集规模增长。

原文摘要 · Abstract (English)

Online reinforcement learning (RL) with sparse rewards poses a challenge partly because of the lack of feedback on states leading to the goal. Furthermore, expert offline data with reward signal is rarely available to provide this feedback and bootstrap online learning. How can we guide online agents to the right solution without this on-task data? Reward shaping offers a solution by providing fine-grained signal to nudge the policy towards the optimal solution. However, reward shaping often requires domain knowledge to hand-engineer heuristics for a specific goal. To enable more general and inexpensive guidance, we propose and analyze a data-driven methodology that automatically guides RL by learning from widely available video data such as Internet recordings, off-task demonstrations, task failures, and undirected environment interaction. By learning a model of optimal goal-conditioned value from diverse passive data, we open the floor to scaling up and using various data sources to model general goal-reaching behaviors relevant to guiding online RL. Specifically, we use intent-conditioned value functions to learn from diverse videos and incorporate these goal-conditioned values into the reward. Our experiments show that video-trained value functions work well with a variety of data sources, exhibit positive transfer from human video pre-training, can generalize to unseen goals, and scale with dataset size.

强化学习视频引导奖励塑形泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。