arXiv:2504.08772cs.LGcs.AI2025-04被引 3

用大模型自动生成离线强化学习奖励,省去人工设计

Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning

  • 用视觉语言大模型分析离线数据自动生成奖励信号
  • 在长序列任务中提升泛化能力,配合稀疏奖励显著增效
  • 适合缺乏标注数据或人力成本高的智能决策场景

在离线强化学习中,从固定数据集学习为那些难以进行实时交互的领域提供了可行方案。然而,为离线数据设计密集奖励信号需要大量人工投入和领域知识。基于人类反馈的强化学习(RLHF)虽为替代方案,但因需人机协同仍成本高昂,促使人们关注自动化奖励生成模型。为此,我们提出基于大视觉语言模型的奖励生成方法(RG-VLM),利用LVLM的推理能力,无需人工参与即可从离线数据中生成奖励信号。RG-VLM在长时序任务中提升了泛化性能,并能与稀疏奖励信号无缝融合,显著改善任务表现,展现出作为辅助奖励信号的巨大潜力。

原文摘要 · Abstract (English)

In offline reinforcement learning (RL), learning from fixed datasets presents a promising solution for domains where real-time interaction with the environment is expensive or risky. However, designing dense reward signals for offline dataset requires significant human effort and domain expertise. Reinforcement learning with human feedback (RLHF) has emerged as an alternative, but it remains costly due to the human-in-the-loop process, prompting interest in automated reward generation models. To address this, we propose Reward Generation via Large Vision-Language Models (RG-VLM), which leverages the reasoning capabilities of LVLMs to generate rewards from offline data without human involvement. RG-VLM improves generalization in long-horizon tasks and can be seamlessly integrated with the sparse reward signals to enhance task performance, demonstrating its potential as an auxiliary reward signal.

强化学习大模型奖励生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。