arXiv:2604.12320cs.CVcs.AI2026-04

构建首个电竞第一视角视频问答基准,评估模型在快节奏虚拟环境中的感知与推理能力。

EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports

论文配图:EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
图 1 · 摘自论文原文
  • 基于三款FPS游戏职业比赛,构建1745对高质量问答数据。
  • 模型最佳表现仅71.58%,在战术推理和微观操作上明显不足。
  • 揭示视觉感知与战略理解间的差距,助力电竞AI优化。

尽管视频大语言模型在慢速、现实的第一人称视频中表现出色,但在高速、信息密集的虚拟环境中其能力仍待深入探索。现有基准集中于日常活动,缺乏对虚拟场景中快速规则推理的严格测评。为此,我们提出EgoEsportsQA,首个面向电竞知识的感知与推理视频问答基准。通过可扩展的六阶段流程,从三款第一人称射击游戏的职业比赛中收集1,745组高质量问答对。问题采用二维解耦分类体系:认知能力维度包含11个子任务(涵盖感知与推理层级),电竞知识维度包含6个子任务。对主流Video-LLM的全面评估显示,当前模型性能未达理想水平,最佳模型仅71.58%。结果揭示显著差距:模型在基础视觉感知优于深层战术推理,在宏观进程理解强于微观操作把握。大量消融实验表明当前Video-LLM架构存在内在缺陷。进一步分析表明,该数据集不仅揭示真实与虚拟第一人称领域的关联,也为下游电竞应用优化提供指导,推动Video-LLM在各类第一人称环境中的未来发展。

原文摘要 · Abstract (English)

While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual environments remain under-explored. Existing benchmarks focus on daily activities, yet lack a rigorous testbed for evaluating fast, rule-bound reasoning in virtual scenarios. To fill this gap, we introduce EgoEsportsQA, a pioneering video question-answering (QA) benchmark for grounding perception and reasoning in expert esports knowledge. We curate 1,745 high-quality QA pairs from professional matches across 3 first-person shooter games via a scalable six-stage pipeline. These questions are structured into a two-dimensional decoupled taxonomy: 11 sub-tasks in the cognitive capability dimension (covering perception and reasoning levels) and 6 sub-tasks in the esports knowledge dimension. Comprehensive evaluations of state-of-the-art Video-LLMs reveal that current models still fail to achieve satisfactory performance, with the best model only 71.58%. The results expose notable gaps across both axes: models exhibit stronger capabilities in basic visual perception than in deep tactical reasoning, and they grasp overall macro-progression better than fine-grained micro-operations. Extensive ablation experiments demonstrate the intrinsic weaknesses of current Video-LLM architectures. Further analysis suggests that our dataset not only reveals the connections between real-world and virtual egocentric domains, but also offers guidance for optimizing downstream esports applications, thereby fostering the future advancement of Video-LLMs in various egocentric environments.

视频问答电竞AI感知推理第一视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。