让多个角色在同一个世界中同步生成视频,保持视角一致。
Prisma-World: Camera-Controllable Multi-Agent Video World Model

- 用联合去噪机制统一多视角生成,确保场景一致性。
- 支持灵活角色数和相机控制,跨视角一致率显著提升。
- 适合虚拟拍摄、多智能体仿真等需要全局空间感知的场景。
视频世界模型在生成可控视觉体验方面进展迅速,但多数仍基于单一观察者视角。扩展至多智能体时的核心挑战是:若各智能体未来状态独立生成,重叠视角可能呈现不同版本的同一场景,导致物体、布局和外观不一致。传统相机条件仅控制单条轨迹,未显式关联应共享场景几何的视图生成。我们提出 Prisma-World,一种可相机控制的多智能体世界模型,将多智能体生成建模为联合几何感知去噪过程以实现跨视角一致性。该模型在单一全注意力序列中处理所有智能体视频,采用多智能体 RoPE 区分身份并保持同步时间坐标,并通过注入相对相机几何信息引导重叠视角聚焦共享场景证据。为进一步强化多视图一致性与全局空间感知,引入重叠衰减课程训练范式与小地图条件结构引导。为支持多智能体模型训练与评估,我们构建 PrismaDataset,一个大规模 UE5 数据集,涵盖多样场景的全景采集、可组合的多智能体视图组(支持灵活智能体数量与复杂相机轨迹)及精确的相机/动作标注。实验表明,单一 Prisma-World 模型可生成高保真多智能体视频,具备灵活智能体数、相机可控性、改进的跨视角一致性与小地图引导下的空间定位能力。
原文摘要 · Abstract (English)
Video world models have made rapid progress in generating controllable visual experiences, but most of them still simulate the world from a single observer. Extending such models to multiple agents raises a central challenge: if each agent's future state is generated independently, overlapping views may instantiate different versions of the same scene, leading to inconsistent objects, layouts, and appearances across agents. Conventional camera conditioning controls individual trajectories, but it does not explicitly couple the generation of views that should agree under shared scene geometry. We introduce Prisma-World, a camera-controllable multi-agent world model that formulates multi-agent generation as a joint geometry-aware denoising process for cross-view consistency. Prisma-World processes all agent videos within one full-attention sequence, uses a multi-agent RoPE design to distinguish agent identities while preserving synchronized temporal coordinates, and injects relative camera geometry into attention to bias overlapping viewpoints toward shared scene evidence. To further strengthen multi-view consistency and enhance global spatial perception, we augment our framework with an overlap-decaying curriculum training paradigm alongside minimap-conditioned structural guidance. To facilitate the training and evaluation of multi-agent models, we introduce PrismaDataset, a large-scale UE5 dataset with panoramic acquisition across diverse scenes, composable multi-agent view groups with flexible agent counts and complex camera trajectories, and precise camera/action annotations for consistency training and evaluation. Experiments show that a single Prisma-World model can generate high-fidelity multi-agent videos with flexible agent numbers, camera controllability, improved cross-view consistency, and spatial grounding under minimap guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。