arXiv:2502.01616cs.LG2025-02被引 5

用视觉语言模型减少人类标注,让强化学习更高效

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning

  • 用VLM生成初始偏好标签,只对不确定样本请人标注
  • 实验显示只需一半标注量,效果仍达顶尖水平
  • 适配后的VLM能跨任务迁移,进一步降低标注需求

基于偏好的强化学习(RL)是实现策略与人类意图对齐的有前景方法,但常受限于高昂的人类反馈成本。本文提出PrefVLM框架,将视觉语言模型(VLM)与选择性人类反馈结合,显著降低标注需求的同时保持性能。方法利用VLM生成初始偏好标签,并筛选出置信度低的样本进行针对性人工标注;同时采用自监督逆动力学损失微调VLM,增强其与演化策略的一致性。在Meta-World操作任务上的实验表明,PrefVLM在使用最多减少2倍人类标注的情况下,成功率达或优于现有最优方法。此外,经适配的VLM可实现任务间高效知识迁移,进一步减少反馈需求。结果表明,结合VLM与选择性人类监督,能有效提升偏好式RL的可扩展性与实用性。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (RL) offers a promising approach for aligning policies with human intent but is often constrained by the high cost of human feedback. In this work, we introduce PrefVLM, a framework that integrates Vision-Language Models (VLMs) with selective human feedback to significantly reduce annotation requirements while maintaining performance. Our method leverages VLMs to generate initial preference labels, which are then filtered to identify uncertain cases for targeted human annotation. Additionally, we adapt VLMs using a self-supervised inverse dynamics loss to improve alignment with evolving policies. Experiments on Meta-World manipulation tasks demonstrate that PrefVLM achieves comparable or superior success rates to state-of-the-art methods while using up to 2 x fewer human annotations. Furthermore, we show that adapted VLMs enable efficient knowledge transfer across tasks, further minimizing feedback needs. Our results highlight the potential of combining VLMs with selective human supervision to make preference-based RL more scalable and practical.

强化学习视觉语言模型偏好学习少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。