arXiv:2608.26200cs.AIcs.CV2026-08

首个支持原生游戏操作的视频游戏世界-动作模型,能生成视觉和控制指令。

GameWAM: A World Action Model for Video Games

论文配图:GameWAM: A World Action Model for Video Games
图 1 · 摘自论文原文
  • 并行生成视觉与键盘鼠标动作,用块因果条件与流匹配建模
  • 任务成功率媲美现有模型,执行动作减少30%以上
  • 适合研究生成式游戏智能、交互控制与长时序决策

现代视频游戏融合第一人称感知、快速视觉变化、持续世界状态及异构原生控制。现有游戏智能体直接将视觉与任务上下文映射为动作,缺乏显式世界动态建模;交互式游戏世界模型可预测视觉未来,但无法作为任务策略。世界-动作模型(WAM)统一二者目标,但在视频游戏动态与开放交互下仍少被探索。本文提出GameWAM,据知是首个支持原生闭环游戏与GUI控制的WAM。GameWAM通过块因果条件与流匹配,联合生成未来视觉观测与可执行的键盘-鼠标轨迹。为支持联合学习,构建同步的游戏进程与GUI轨迹数据。针对异构原生控制,GameWAM在每步预测游戏/GUI模式,使用模式特定分布与连续动作归一化生成动作。针对长时交互,块循环控制预测超出承诺周期,仅执行短动作前缀并从新观测重规划,同时细粒度循环内上下文与分层跨周期历史保持时间连续性。实验显示任务成功率优于或相当,执行原生动作数减少超30%。进一步发现低频动作源印刻(LASI)现象:采样动作源的低频成分在固定条件下系统性引导粗粒度相机运动,揭示生成控制中的源敏感性失效模式。

原文摘要 · Abstract (English)

Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied actions but do not serve as task policies. World-Action Models (WAMs) unify these objectives, but remain largely unexplored under the dynamics and open-ended interaction of video games. We introduce GameWAM, to our knowledge the first WAM for native closed-loop gameplay and GUI control. GameWAM jointly generates future visual observations and executable keyboard-mouse trajectories through parallel visual and action generative processes with block-causal conditioning and flow matching. To support joint world-action learning, we construct synchronized gameplay and GUI trajectories. To handle heterogeneous native control, GameWAM predicts a gameplay/GUI mode at each action step and generates actions with mode-specific prediction distributions and continuous-action normalization. For long-horizon interaction, block-cycle control predicts beyond the committed horizon, executes only a short action prefix, and replans from new observations, while fine-grained within-cycle context and hierarchical cross-cycle history preserve temporal continuity. Experiments demonstrate competitive task success with fewer executed native actions than the compared agents. We further uncover Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control. Project page is available at https://yunncheng.github.io/GameWAM/.

游戏智能生成模型动作控制世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。