arXiv:2607.00483cs.RO2026-07中稿 · IJCAI

用视觉语言模型同时生成绝对与相对奖励,提升强化学习任务表现

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning

论文配图:VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning
图 1 · 摘自论文原文
  • 利用视觉语言模型理解任务目标并生成奖励标签
  • 结合状态评估与对比监督,提升奖励信号稳定性与鲁棒性
  • 在Minecraft等复杂场景中显著优于已有方法

强化学习中的有效奖励函数设计仍是重大挑战,尤其在开放环境里任务目标抽象难量化。本文提出VLM-AR3L框架,利用视觉语言模型(VLM)提供绝对与相对奖励。该框架将智能体的视觉观测置于自然语言任务目标的上下文中,从VLM生成的偏好标签中学习两种奖励:绝对奖励模型预测单个状态的标量评估值,相对奖励模型通过比较连续观测判断任务进展或退步。二者融合兼顾状态评估的稳定性与对比监督的鲁棒性。在经典控制、操作任务及开放世界具身任务多个基准上进行评估,重点关注Minecraft这一视觉复杂且需长时决策的环境。实验表明,VLM-AR3L持续优于现有基于VLM的奖励学习方法。

原文摘要 · Abstract (English)

Designing effective reward functions remains a major challenge in reinforcement learning (RL), particularly in open-ended environments where task goals are abstract and difficult to quantify. In this work, we present VLM-AR3L, a framework that leverages Vision-Language Models (VLMs) to provide both absolute and relative rewards for RL. VLM-AR3L interprets an agent's visual observations in the context of a natural language task goal, and learns both absolute and relative rewards from VLM-generated preference labels. The absolute reward model predicts scalar evaluations for individual states, while the relative reward model compares consecutive observations to infer progress or regression toward the task goal. Their integration combines the stability of state-based evaluation with the robustness of comparative supervision. We evaluate VLM-AR3L across benchmarks spanning classic control, manipulation, and open-world embodied tasks, with a particular focus on Minecraft given its visual complexity and long-horizon decision-making requirements. Experimental results show that VLM-AR3L consistently outperforms prior VLM-based reward learning methods.

强化学习视觉语言模型奖励设计Minecraft

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。