arXiv:2604.18564cs.CV2026-04被引 10

多智能体多视角视频世界模型,可高效模拟复杂交互场景

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

论文配图:MultiWorld: Scalable Multi-Agent Multi-View Video World Models
图 1 · 摘自论文原文
  • 设计多智能体条件模块与全局状态编码器,实现多智能体精准控制
  • 在多人游戏和机器人任务中,视频保真度与跨视角一致性显著优于基线
  • 支持灵活扩展智能体与视角数量,多视图并行生成提升效率

视频世界模型在响应用户或智能体动作时,对环境动态的模拟已取得显著进展。它们被建模为以历史帧和当前动作作为输入的动作条件视频生成模型,用于预测未来帧。然而,大多数现有方法局限于单智能体场景,难以捕捉真实多智能体系统中的复杂交互。本文提出 extbf{MultiWorld},一个统一的多智能体多视角世界建模框架,可在保持多视角一致性的同时实现对多个智能体的精确控制。我们引入多智能体条件模块以实现精细的多智能体可控性,并设计全局状态编码器以确保不同视角下观测的一致性。MultiWorld 支持灵活扩展智能体与视角数量,且能并行合成多视角内容,具备高效率。在多人游戏环境和多机器人操作任务上的实验表明,MultiWorld 在视频保真度、动作跟随能力以及多视角一致性方面均超越基线方法。

原文摘要 · Abstract (English)

Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as action-conditioned video generation models that take historical frames and current actions as input to predict future frames. Yet, most existing approaches are limited to single-agent scenarios and fail to capture the complex interactions inherent in real-world multi-agent systems. We present \textbf{MultiWorld}, a unified framework for multi-agent multi-view world modeling that enables accurate control of multiple agents while maintaining multi-view consistency. We introduce the Multi-Agent Condition Module to achieve precise multi-agent controllability, and the Global State Encoder to ensure coherent observations across different views. MultiWorld supports flexible scaling of agent and view counts, and synthesizes different views in parallel for high efficiency. Experiments on multi-player game environments and multi-robot manipulation tasks demonstrate that MultiWorld outperforms baselines in video fidelity, action-following ability, and multi-view consistency. Project page: https://multi-world.github.io/

多智能体视频生成世界模型多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。