大模型能准确识别游戏中的玩家投入度吗?
Can Large Language Models Capture Video Game Engagement?
- 用文本+画面帧提示大模型预测游戏视频情感变化
- 4800次实验发现模型表现接近人类但仍有差距
- 适合研究自动情绪标注与人机交互的学者
为探究预训练大语言模型(LLMs)在无额外训练情况下,能否成功检测观看视频时的人类情感,本文首次全面评估了主流LLMs在多模态提示(文本序列+视频帧)下对连续情感标注的预测能力。基于GameVibe数据集,我们测试了20款第一人称射击游戏中80分钟已标注的游戏视频片段中,模型对游戏内投入度变化的识别能力。共开展超过4,800次实验,考察了模型架构、规模、输入模态、提示策略及真实标签处理方法对预测效果的影响。结果表明,尽管LLMs在多个领域表现出类人性能,且优于传统机器学习基线,但仍普遍落后于人类提供的连续体验标注。我们分析了跨游戏表现波动的原因,指出了模型超出预期的情况,并为未来通过LLMs实现自动化情感标注提供了发展路线图。
原文摘要 · Abstract (English)
Can out-of-the-box pretrained Large Language Models (LLMs) detect human affect successfully when observing a video? To address this question, for the first time, we evaluate comprehensively the capacity of popular LLMs for successfully predicting continuous affect annotations of videos when prompted by a sequence of text and video frames in a multimodal fashion. In this paper, we test LLMs' ability to correctly label changes of in-game engagement in 80 minutes of annotated videogame footage from 20 first-person shooter games of the GameVibe corpus. We run over 4,800 experiments to investigate the impact of LLM architecture, model size, input modality, prompting strategy, and ground truth processing method on engagement prediction. Our findings suggest that while LLMs rightfully claim human-like performance across multiple domains and able to outperform traditional machine learning baselines, they generally fall behind continuous experience annotations provided by humans. We examine some of the underlying causes for a fluctuating performance across games, highlight the cases where LLMs exceed expectations, and draw a roadmap for the further exploration of automated emotion labelling via LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。