arXiv:2411.13211cs.CVcs.LG2024-11

测试视觉语言模型对序列任务的理解能力,发现多数模型表现不佳。

ViSTa Dataset: Do vision-language models understand sequential tasks?

  • 构建了包含4000+视频的分层序列任务数据集ViSTa
  • GPT-4o是唯一达到非平凡表现的模型,其他模型在复杂任务上失败
  • 适用于研究模型对动态行为理解能力的研究者

将视觉语言模型(VLMs)作为强化学习中的奖励模型,有望降低训练成本并提升安全性。目前,这类模型仅用于目标导向任务,即仅通过最终状态评分。本文探索其在无法仅由终态评估的任务中的潜力。为此,提出ViSTa数据集,用于评估基于视觉的序列任务理解能力。ViSTa包含超过4,000个视频,涵盖虚拟家居、Minecraft及真实世界环境,具有从单步任务逐步组合成更复杂序列任务的分层结构,可精细衡量VLM在不同复杂度下的判断能力。我们使用ViSTa评估CLIP、ViCLIP和GPT-4o等先进VLM,发现尽管它们在物体识别上表现良好,但在序列任务理解上普遍失败,仅有GPT-4o展现出非平凡性能。

原文摘要 · Abstract (English)

Using vision-language models (VLMs) as reward models in reinforcement learning holds promise for reducing costs and improving safety. So far, VLM reward models have only been used for goal-oriented tasks, where the agent must reach a particular final outcome. We explore VLMs' potential to supervise tasks that cannot be scored by the final state alone. To this end, we introduce ViSTa, a dataset for evaluating Vision-based understanding of Sequential Tasks. ViSTa comprises over 4,000 videos with step-by-step descriptions in virtual home, Minecraft, and real-world environments. Its novel hierarchical structure -- basic single-step tasks composed into more and more complex sequential tasks -- allows a fine-grained understanding of how well VLMs can judge tasks with varying complexity. To illustrate this, we use ViSTa to evaluate state-of-the-art VLMs, including CLIP, ViCLIP, and GPT-4o. We find that, while they are all good at object recognition, they fail to understand sequential tasks, with only GPT-4o achieving non-trivial performance.

视觉语言模型序列理解强化学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。