arXiv:2603.01694cs.CVcs.AI2026-03

用多视角视频提升强化学习奖励设计,让机器人动作更自然。

MVR: Multi-view Video Reward Shaping for Reinforcement Learning

  • 从多个视角拍摄视频,用视觉语言模型评估状态相关性。
  • 动态调整奖励权重,达成目标后自动降低模型引导强度。
  • 在人形机器人行走和操作任务中显著提升训练效果。

奖励设计对强化学习解决复杂任务至关重要。近期研究利用视觉语言模型(VLM)生成的图像-文本相似度来增强任务视觉反馈奖励,但通常线性叠加且无显式塑形,可能改变最优策略。此外,这些方法多依赖单张静态图像,难以处理涉及复杂动态行为的任务,且单一视角易遮挡关键动作信息。为此,本文提出多视角视频奖励塑形(MVR)框架,通过多视角视频建模状态与目标任务的相关性。MVR利用冻结的预训练VLM计算视频-文本相似度,学习状态相关性函数,缓解图像方法对特定静态姿态的偏差。同时引入状态相关奖励塑形机制,融合任务特定奖励与VLM指导,一旦达成期望运动模式,自动减弱VLM引导影响。我们在HumanoidBench的人形机器人行走任务和MetaWorld的操作任务上进行了大量实验,验证了该框架的有效性,并通过消融实验确认了设计选择的合理性。

原文摘要 · Abstract (English)

Reward design is of great importance for solving complex tasks with reinforcement learning. Recent studies have explored using image-text similarity produced by vision-language models (VLMs) to augment rewards of a task with visual feedback. A common practice linearly adds VLM scores to task or success rewards without explicit shaping, potentially altering the optimal policy. Moreover, such approaches, often relying on single static images, struggle with tasks whose desired behavior involves complex, dynamic motions spanning multiple visually different states. Furthermore, single viewpoints can occlude critical aspects of an agent's behavior. To address these issues, this paper presents Multi-View Video Reward Shaping (MVR), a framework that models the relevance of states regarding the target task using videos captured from multiple viewpoints. MVR leverages video-text similarity from a frozen pre-trained VLM to learn a state relevance function that mitigates the bias towards specific static poses inherent in image-based methods. Additionally, we introduce a state-dependent reward shaping formulation that integrates task-specific rewards and VLM-based guidance, automatically reducing the influence of VLM guidance once the desired motion pattern is achieved. We confirm the efficacy of the proposed framework with extensive experiments on challenging humanoid locomotion tasks from HumanoidBench and manipulation tasks from MetaWorld, verifying the design choices through ablation studies.

强化学习视频理解奖励设计多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。