arXiv:2505.06482cs.LGcs.AI2025-05ICML被引 1

用视频数据构建世界模型,提升离线强化学习性能

Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach

  • 从海量视频构建可交互的世界模型,弥补无环境交互缺陷
  • 在机器人、自动驾驶等任务中性能提升超100%
  • 适合需要安全高效训练的视觉控制场景

离线强化学习利用静态数据集优化策略,避免了真实世界探索的风险与成本。然而,由于缺乏环境交互,其常面临次优行为和价值估计不准的问题。本文提出基于模型的视频增强离线强化学习(VeoRL),通过利用在线可获取的多样化未标注视频数据构建一个交互式世界模型。借助基于模型的行为引导,该方法将自然视频中的控制策略与物理动态常识迁移至目标领域的强化学习智能体。VeoRL在机器人操作、自动驾驶及开放世界视频游戏等视觉控制任务中均取得显著性能提升,部分任务表现提升超过100%。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) enables policy optimization using static datasets, avoiding the risks and costs of extensive real-world exploration. However, it struggles with suboptimal offline behaviors and inaccurate value estimation due to the lack of environmental interaction. We present Video-Enhanced Offline RL (VeoRL), a model-based method that constructs an interactive world model from diverse, unlabeled video data readily available online. Leveraging model-based behavior guidance, our approach transfers commonsense knowledge of control policy and physical dynamics from natural videos to the RL agent within the target domain. VeoRL achieves substantial performance gains (over 100% in some cases) across visual control tasks in robotic manipulation, autonomous driving, and open-world video games.

离线强化学习视频生成世界模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。