arXiv:2512.20675cs.LGcs.AI2025-12被引 2

简单对比损失比复杂方法更有效,揭示了视觉语言奖励模型的关键瓶颈。

Revisiting the Learning Objectives of Vision-Language Reward Models

  • 统一框架下对比不同学习目标,控制数据、模型和评估环境一致。
  • 三元组损失在Meta-World任务上超越当前最优方法,准确率更高。
  • 适合研究多模态强化学习与奖励建模的开发者参考。

在具身智能中,学习通用奖励函数是一项核心挑战。近期工作利用对比视觉语言模型(VLM)获得无需人工标注的密集、领域无关奖励。这些方法通过日益复杂的训练目标将VLM转化为奖励模型,但由于训练数据、架构和评估设置的差异,难以进行有意义的比较。本文通过统一框架,使用相同主干网络、微调数据和评估环境,评估了近期基于VLM的奖励模型。在Meta-World任务中,我们以与真实奖励的一致性和与专家进展的相关性为指标衡量建模精度。结果表明,简单的三元组损失优于现有最先进方法,暗示近期方法的提升更多源于数据和架构差异,而非学习目标本身。

原文摘要 · Abstract (English)

Learning generalizable reward functions is a core challenge in embodied intelligence. Recent work leverages contrastive vision language models (VLMs) to obtain dense, domain-agnostic rewards without human supervision. These methods adapt VLMs into reward models through increasingly complex learning objectives, yet meaningful comparison remains difficult due to differences in training data, architectures, and evaluation settings. In this work, we isolate the impact of the learning objective by evaluating recent VLM-based reward models under a unified framework with identical backbones, finetuning data, and evaluation environments. Using Meta-World tasks, we assess modeling accuracy by measuring consistency with ground truth reward and correlation with expert progress. Remarkably, we show that a simple triplet loss outperforms state-of-the-art methods, suggesting that much of the improvements in recent approaches could be attributed to differences in data and architectures.

视觉语言奖励模型强化学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。