arXiv:2604.02330cs.CVcs.AI2026-04被引 1

让视频生成模型能同时控制多个角色,且动作与角色绑定更准确。

ActionParty: Multi-Subject Action Binding in Generative Video Games

论文配图:ActionParty: Multi-Subject Action Binding in Generative Video Games
图 1 · 摘自论文原文
  • 用状态令牌持久记录每个角色的动态,实现多角色独立控制。
  • 在46个环境中成功同步控制7名玩家,动作跟随准确率显著提升。
  • 适合开发多人互动视频游戏或复杂场景仿真系统的人参考。

视频扩散模型的发展使得“世界模型”得以模拟交互式环境,但现有模型多局限于单智能体设置,难以同时控制场景中的多个智能体。本文针对视频扩散模型中动作与主体绑定困难的核心问题,提出ActionParty——一个可控制多主体的生成式视频游戏世界模型。该模型引入主体状态令牌(即持续捕捉场景中每个主体状态的潜在变量),通过空间偏置机制联合建模状态令牌与视频潜在表示,将全局画面渲染与个体动作驱动的主体更新解耦。在Melting Pot基准上评估显示,ActionParty是首个能在46种不同环境中同步控制最多七名玩家的视频世界模型,结果表明其在动作跟随准确性和身份一致性方面均有显著提升,并能稳健地通过复杂交互实现主体的自回归追踪。

原文摘要 · Abstract (English)

Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However, these models are largely restricted to single-agent settings, failing to control multiple agents simultaneously in a scene. In this work, we tackle a fundamental issue of action binding in existing video diffusion models, which struggle to associate specific actions with their corresponding subjects. For this purpose, we propose ActionParty, an action controllable multi-subject world model for generative video games. It introduces subject state tokens, i.e. latent variables that persistently capture the state of each subject in the scene. By jointly modeling state tokens and video latents with a spatial biasing mechanism, we disentangle global video frame rendering from individual action-controlled subject updates. We evaluate ActionParty on the Melting Pot benchmark, demonstrating the first video world model capable of controlling up to seven players simultaneously across 46 diverse environments. Our results show significant improvements in action-following accuracy and identity consistency, while enabling robust autoregressive tracking of subjects through complex interactions.

视频生成多主体动作控制世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。