arXiv:2510.19150cs.CVcs.AI2025-10被引 2

通过队友视角对比学习,让单个玩家看清团队战术全局。

X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning

  • 用多视角同步视频对比学习,对齐队友第一视角
  • 仅凭一人视角就可预测队友和对手位置,准确率显著提升
  • 适合研究电竞、人机协同与多智能体决策的学者

人类团队战术源于个体视角及对队友意图的预判、理解与适应。尽管视频理解进展提升了体育赛事中团队互动建模能力,但现有工作多依赖第三人称直播画面,忽视了多智能体学习中同步的第一人称特性。我们提出X-Ego-CS基准数据集,包含45场职业级《反恐精英2》比赛的124小时游戏录像,提供同步的跨第一人称视角视频流与状态-动作轨迹。基于此,我们设计交叉第一人称对比学习(CECL),对齐队友的第一人称视觉流,使个体视角具备团队级战术态势感知能力。在队友/对手位置预测任务上,使用先进视频编码器验证了其有效性,证明仅凭单一第一人称视角即可准确推断双方位置。X-Ego-CS与CECL共同构建了电子竞技中跨第一人称多智能体评估的基础。更广泛地,本工作将游戏理解作为多智能体建模与战术学习的测试平台,对时空推理与人机协同具有深远意义。代码与数据集见https://github.com/HATS-ICT/x-ego。

原文摘要 · Abstract (English)

Human team tactics emerge from each player's individual perspective and their ability to anticipate, interpret, and adapt to teammates' intentions. While advances in video understanding have improved the modeling of team interactions in sports, most existing work relies on third-person broadcast views and overlooks the synchronous, egocentric nature of multi-agent learning. We introduce X-Ego-CS, a benchmark dataset consisting of 124 hours of gameplay footage from 45 professional-level matches of the popular e-sports game Counter-Strike 2, designed to facilitate research on multi-agent decision-making in complex 3D environments. X-Ego-CS provides cross-egocentric video streams that synchronously capture all players' first-person perspectives along with state-action trajectories. Building on this resource, we propose Cross-Ego Contrastive Learning (CECL), which aligns teammates' egocentric visual streams to foster team-level tactical situational awareness from an individual's perspective. We evaluate CECL on a teammate-opponent location prediction task, demonstrating its effectiveness in enhancing an agent's ability to infer both teammate and opponent positions from a single first-person view using state-of-the-art video encoders. Together, X-Ego-CS and CECL establish a foundation for cross-egocentric multi-agent benchmarking in esports. More broadly, our work positions gameplay understanding as a testbed for multi-agent modeling and tactical learning, with implications for spatiotemporal reasoning and human-AI teaming in both virtual and real-world domains. Code and dataset are available at https://github.com/HATS-ICT/x-ego.

多智能体电竞态势感知对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。