arXiv:2603.18480cs.CVcs.AI2026-03被引 2

测试视觉语言模型能否从游戏画面推断玩家投入度,发现效果有限。

Do Vision Language Models Understand Human Engagement in Games?

  • 用九款射击游戏数据,对比六种提示策略评估模型表现
  • 零样本预测普遍较弱,多数不如简单基线,成对变化预测更难
  • 理论引导提示反使模型依赖表面特征,暴露理解短板

从游戏视频中推断人类投入度对游戏设计与玩家体验研究至关重要,但当前视觉-语言模型(VLMs)是否仅凭视觉线索就能推断这种潜在心理状态仍不明确。我们基于九款第一人称射击游戏的GameVibe少样本数据集,评估三种VLM在六种提示策略下的表现,包括零样本预测、基于心流、游戏流、自我决定理论及MDA框架的理论引导提示,以及检索增强提示。考察点预测与连续窗口间投入度变化的成对预测。结果表明,零样本预测普遍表现不佳,常不及每游戏的多数类基线;检索增强提示在部分场景下提升点预测表现,但成对预测始终困难。单纯使用理论引导提示并未可靠提升性能,反而可能强化表层捷径。这些发现表明当前VLM存在感知-理解鸿沟:虽能识别可见游戏线索,却难以在跨游戏场景中稳健推断人类投入度。

原文摘要 · Abstract (English)

Inferring human engagement from gameplay video is important for game design and player-experience research, yet it remains unclear whether vision--language models (VLMs) can infer such latent psychological states from visual cues alone. Using the GameVibe Few-Shot dataset across nine first-person shooter games, we evaluate three VLMs under six prompting strategies, including zero-shot prediction, theory-guided prompts grounded in Flow, GameFlow, Self-Determination Theory, and MDA, and retrieval-augmented prompting. We consider both pointwise engagement prediction and pairwise prediction of engagement change between consecutive windows. Results show that zero-shot VLM predictions are generally weak and often fail to outperform simple per-game majority-class baselines. Memory- or retrieval-augmented prompting improves pointwise prediction in some settings, whereas pairwise prediction remains consistently difficult across strategies. Theory-guided prompting alone does not reliably help and can instead reinforce surface-level shortcuts. These findings suggest a perception--understanding gap in current VLMs: although they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games.

视觉语言模型游戏行为分析心理状态推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。