arXiv:2506.12822cs.LGcs.RO2025-06ICML被引 18

用AI生成评分提升强化学习效率,减少人工标注。

Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models

  • 让大视觉语言模型直接打分单条轨迹,比对更高效
  • 在低层和高层控制任务中均显著超越现有方法
  • 适合想降低人工反馈成本的研究者

强化学习中的奖励函数设计仍是核心挑战,通常需大量人力与领域知识。尽管基于人类反馈的强化学习已成功对齐人类意图,但高质量反馈获取成本高、难以扩展。近年来,基础模型的发展提供了新路径——利用AI生成反馈以减少对人类监督的依赖。本文提出ERL-VLM,一种增强型评分式强化学习方法,可有效从AI反馈中学习奖励函数。不同于以往依赖成对比较的方法,ERL-VLM通过大视觉语言模型(VLM)对单条轨迹给出绝对评分,实现更丰富的反馈并提升样本效率。此外,我们引入关键改进,缓解数据不平衡与标签噪声导致的训练不稳定性。在低层与高层控制任务上的广泛实验表明,ERL-VLM显著优于现有基于VLM的奖励生成方法。结果证明了AI反馈在实现低人工干预下规模化强化学习的潜力,为更自主高效的奖励学习铺平道路。

原文摘要 · Abstract (English)

Designing effective reward functions remains a fundamental challenge in reinforcement learning (RL), as it often requires extensive human effort and domain expertise. While RL from human feedback has been successful in aligning agents with human intent, acquiring high-quality feedback is costly and labor-intensive, limiting its scalability. Recent advancements in foundation models present a promising alternative--leveraging AI-generated feedback to reduce reliance on human supervision in reward learning. Building on this paradigm, we introduce ERL-VLM, an enhanced rating-based RL method that effectively learns reward functions from AI feedback. Unlike prior methods that rely on pairwise comparisons, ERL-VLM queries large vision-language models (VLMs) for absolute ratings of individual trajectories, enabling more expressive feedback and improved sample efficiency. Additionally, we propose key enhancements to rating-based RL, addressing instability issues caused by data imbalance and noisy labels. Through extensive experiments across both low-level and high-level control tasks, we demonstrate that ERL-VLM significantly outperforms existing VLM-based reward generation methods. Our results demonstrate the potential of AI feedback for scaling RL with minimal human intervention, paving the way for more autonomous and efficient reward learning.

强化学习AI反馈视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。