arXiv:2511.23429cs.CV2025-11被引 46

用自然语言控制游戏世界,让生成内容更灵活互动。

Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model

  • 通过自然语言指令实现对游戏世界的细粒度控制
  • 生成时序连贯且因果一致的交互式游戏视频
  • 适合游戏开发、AI创作与交互式内容生成研究者

生成式世界模型近年推动了开放世界游戏环境的创建,从静态场景合成发展为动态交互模拟。然而,现有方法受限于固定的动作规则和高昂的标注成本,难以建模多样化的游戏内交互与玩家驱动的动态行为。为此,我们提出 Hunyuan-GameCraft-2,一种指令驱动的交互式生成游戏世界新范式。用户可通过自然语言、键盘或鼠标信号控制生成视频内容,突破传统固定输入的限制。我们首次形式化定义了交互式视频数据,并建立自动化流程,将大规模无结构文本-视频对转换为因果对齐的交互数据集。基于14B参数的图像到视频MoE基础模型,引入文本驱动的交互注入机制,实现对摄像机运动、角色行为和环境动态的精细调控。我们构建了交互评估基准 InterBench,实验证明模型能准确响应“打开门”“画一个火把”“触发爆炸”等自由形式指令,生成具有时间一致性与因果合理性的交互式游戏视频。

原文摘要 · Abstract (English)

Recent advances in generative world models have enabled remarkable progress in creating open-ended game environments, evolving from static scene synthesis toward dynamic, interactive simulation. However, current approaches remain limited by rigid action schemas and high annotation costs, restricting their ability to model diverse in-game interactions and player-driven dynamics. To address these challenges, we introduce Hunyuan-GameCraft-2, a new paradigm of instruction-driven interaction for generative game world modeling. Instead of relying on fixed keyboard inputs, our model allows users to control game video contents through natural language prompts, keyboard, or mouse signals, enabling flexible and semantically rich interaction within generated worlds. We formally defined the concept of interactive video data and developed an automated process to transform large-scale, unstructured text-video pairs into causally aligned interactive datasets. Built upon a 14B image-to-video Mixture-of-Experts(MoE) foundation model, our model incorporates a text-driven interaction injection mechanism for fine-grained control over camera motion, character behavior, and environment dynamics. We introduce an interaction-focused benchmark, InterBench, to evaluate interaction performance comprehensively. Extensive experiments demonstrate that our model generates temporally coherent and causally grounded interactive game videos that faithfully respond to diverse and free-form user instructions such as "open the door", "draw a torch", or "trigger an explosion".

游戏生成指令控制视频生成交互模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。