arXiv:2603.24329cs.CLcs.AI2026-03ACL被引 2

构建游戏视频理解新基准,评估智能体在3D环境中的多视角决策能力。

GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents

论文配图:GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
图 1 · 摘自论文原文
  • 按1.22标签/秒密度标注多人游戏视频,三元系统分解自我的行为、他者与世界状态。
  • 生成2.4K诊断问答对,模型在时间对齐和角色归属上表现远低于人类。
  • 适合研究具身智能、代理感知与世界建模的学者,支持细粒度错误分析。

多模态大模型正被广泛用作3D环境中自主智能体的感知核心,涵盖机器人到虚拟世界。这些应用要求智能体能感知快速状态变化,准确归因动作主体,并从第一人称视角推理多个智能体的并发行为,而现有基准未能充分评估此类能力。本文提出GameplayQA,一个以智能体为中心的视频理解评估框架。具体而言,我们以1.22标签/秒的密度对多人3D游戏视频进行密集标注,包含时间同步的实时状态、动作与事件描述,采用自我、其他智能体与世界三元结构,适配多智能体环境。基于此标注,构建了2.4K组诊断型问答对,分为三个认知复杂度层级,并设计结构化干扰项分类,支持对模型幻觉位置的精细分析。对前沿多模态大模型的评估显示,其与人类表现存在显著差距,常见失败包括时间与跨视频定位不准、智能体角色归属错误,以及应对高密度决策时的处理不足。我们希望GameplayQA能推动具身智能、代理感知与世界建模交叉领域的研究。

原文摘要 · Abstract (English)

Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct entities, and reason about concurrent multi-agent behaviors from a first-person perspective, capabilities that existing benchmarks do not adequately evaluate. We introduce GameplayQA, a framework for evaluating agentic-centric perception and reasoning through video understanding. Specifically, we densely annotate multiplayer 3D gameplay videos at 1.22 labels/second, with time-synced, concurrent captions of states, actions, and events structured around a triadic system of Self, Other Agents, and the World, a natural decomposition for multi-agent environments. From these annotations, we refined 2.4K diagnostic QA pairs organized into three levels of cognitive complexity, accompanied by a structured distractor taxonomy that enables fine-grained analysis of where models hallucinate. Evaluation of frontier MLLMs reveals a substantial gap from human performance, with common failures in temporal and cross-video grounding, agent-role attribution, and handling the decision density of the game. We hope GameplayQA stimulates future research at the intersection of embodied AI, agentic perception, and world modeling.

多智能体视频理解具身智能测评基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。