arXiv:2411.00769cs.CVcs.AI2024-11ICLR被引 130

首个可交互控制的开放世界游戏视频生成模型

GameGen-X: Interactive Open-world Game Video Generation

论文配图:GameGen-X: Interactive Open-world Game Video Generation
图 1 · 摘自论文原文
  • 基于扩散Transformer架构,支持长序列高质量生成
  • 在百万级视频数据上训练,实现动态角色与场景联动
  • 首次统一角色交互与场景控制,适合游戏开发与模拟应用

我们提出GameGen-X,首个专为生成和交互式控制开放世界游戏视频设计的扩散Transformer模型。该模型通过模拟大量游戏引擎特性,如创新角色、动态环境、复杂动作和多样事件,实现高质量、开放域的游戏视频生成。同时具备交互可控性,可根据当前视频片段预测并调整未来内容,支持游戏玩法模拟。为此,我们从零构建了首个且最大的开放世界游戏视频数据集,包含超百万条来自150多款游戏的多样化游戏视频片段,并附有GPT-4o生成的详细描述。GameGen-X采用两阶段训练:先通过文本到视频生成与视频续写预训练,获得长序列高质量生成能力;再引入InstructNet模块,集成多模态控制信号专家,使模型能根据用户输入调整潜在表示,首次在视频生成中统一角色互动与场景控制。指令微调阶段仅更新InstructNet,冻结预训练基础模型,确保交互能力融入的同时不损失生成多样性与质量。

原文摘要 · Abstract (English)

We introduce GameGen-X, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos. This model facilitates high-quality, open-domain generation by simulating an extensive array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, predicting and altering future content based on the current clip, thus allowing for gameplay simulation. To realize this vision, we first collected and built an Open-World Video Game Dataset from scratch. It is the first and largest dataset for open-world game video generation and control, which comprises over a million diverse gameplay video clips sampling from over 150 games with informative captions from GPT-4o. GameGen-X undergoes a two-stage training process, consisting of foundation model pre-training and instruction tuning. Firstly, the model was pre-trained via text-to-video generation and video continuation, endowing it with the capability for long-sequence, high-quality open-domain game video generation. Further, to achieve interactive controllability, we designed InstructNet to incorporate game-related multi-modal control signal experts. This allows the model to adjust latent representations based on user inputs, unifying character interaction and scene content control for the first time in video generation. During instruction tuning, only the InstructNet is updated while the pre-trained foundation model is frozen, enabling the integration of interactive controllability without loss of diversity and quality of generated video content.

游戏生成扩散模型交互控制视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。