用视觉语言模型为游戏画面自动打奖励标签,提升强化学习训练效率。
VLMs for Videogame Data Annotation

- 用VLM分析游戏帧序列,自动生成奖励信号。
- 模型在赛车游戏上表现不佳,需优化提示词和输出融合。
- 输入长度、分辨率和批量处理影响标注质量与成本。
视觉语言模型(VLMs)和人工智能代理已革新工程师解决现实复杂问题的方式,但在视频游戏中的应用受限于合成场景的极端多样性及与真实物理规律的偏差。本文研究了使用VLM为视频游戏帧序列标注奖励信号的可能性,该任务在条件训练和离线强化学习中具有重要应用价值。结果表明,VLM在赛车类游戏中常无法回答基础问题(其他游戏类型亦然),并提出输出混合与提示优化等应对策略。同时发现,输入序列长度、图像分辨率及问题批处理方式显著影响标注质量与令牌消耗。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization. We also show how input sequence length, resolution, and question batching affect the annotation quality and its token consumption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。