arXiv:2505.15952cs.CVcs.AI2025-05NeurIPS被引 13

构建游戏质检专用评测基准,评估视觉语言模型在真实场景中的表现

VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance

  • 设计涵盖图像/视频的多类游戏质检任务的综合评测体系
  • 首次系统评估VLM在游戏视觉缺陷检测与报告生成中的能力
  • 适合游戏开发自动化、AI质检研究者参考

随着电子游戏成为娱乐产业收入最高的领域,优化开发流程对行业持续发展至关重要。近年来,视觉语言模型(VLMs)在自动化和提升游戏开发各环节方面展现出巨大潜力,尤其是在质量保证(QA)这一高度依赖人力且自动化程度低的环节。为准确评估VLM在游戏质检任务中的表现,并判断其在真实场景下的有效性,亟需标准化评测基准——而现有基准难以满足该领域的特定需求。为此,我们提出VideoGameQA-Bench,一个覆盖视觉单元测试、视觉回归测试、'大海捞针'任务、漏洞检测以及针对多种游戏的图像与视频的缺陷报告生成等广泛游戏质检活动的综合性评测基准。代码与数据已公开:https://asgaardlab.github.io/videogameqa-bench/

原文摘要 · Abstract (English)

With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for the sector's sustained growth. Recent advancements in Vision-Language Models (VLMs) offer considerable potential to automate and enhance various aspects of game development, particularly Quality Assurance (QA), which remains one of the industry's most labor-intensive processes with limited automation options. To accurately evaluate the performance of VLMs in video game QA tasks and determine their effectiveness in handling real-world scenarios, there is a clear need for standardized benchmarks, as existing benchmarks are insufficient to address the specific requirements of this domain. To bridge this gap, we introduce VideoGameQA-Bench, a comprehensive benchmark that covers a wide array of game QA activities, including visual unit testing, visual regression testing, needle-in-a-haystack tasks, glitch detection, and bug report generation for both images and videos of various games. Code and data are available at: https://asgaardlab.github.io/videogameqa-bench/

游戏AI视觉语言模型质量保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。