用单视角视频训练多智能体世界模型,突破拍摄成本与视角同步难题。
MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data

- 通过分解单视角视频中的相机运动与主体轨迹,实现多视角同步
- 跨视图一致性提升37.6%,身份保真度达92.1%
- 适合需要低成本构建多智能体虚拟环境的研究者
视频世界模型是具身智能和元宇宙的基础生成技术,但现有方法仅限于单一智能体从单一视角观察。扩展到多智能体场景面临两大挑战:数据稀缺(通用开放域下协同多视角录制成本过高)和世界状态对齐(独立生成的视频流无法保证共享物理环境与事件在不同视角下一致演化)。为此,我们提出MetaWorld,一种直接从单视角视频扩展至开放域多智能体视频世界模型的新框架。首先,引入单目世界状态展开(MWSU),显式将单视角视频分解为摄像机操作者的自我运动与可见主体的空间轨迹,自然提取共享三维空间内的多智能体同步运动数据,无需多摄像头设置。其次,为实现精确视觉控制,设计了主体感知世界生成器,支持基于个体身份图像的外观驱动模拟。最后,提出世界状态对齐(WSA),在视频DiT每个Transformer层插入帧级跨分支交叉注意力机制,联合同步去噪过程,强制静态几何与动态运动一致性,确保双主视角共享同一物理现实。大量实验表明,MetaWorld在跨视角一致性与身份保真度上均优于现有方法,建立了一种高可扩展、物理驱动的多智能体视频世界建模范式。
原文摘要 · Abstract (English)
Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective. Extending these models to multi-agent settings introduces two critical challenges: data scarcity (coordinated multi-view recordings are prohibitively expensive to collect for general open-domain scenarios) and world state alignment (independently generated video streams cannot ensure that shared physical environments and events evolve consistently across views). To address these challenges, we propose MetaWorld, a novel framework that scales multi-agent video world models to open-domain environments directly from single-view videos. First, we introduce Monocular World-State Unrolling (MWSU) to explicitly decompose monocular footage into the camera operator's ego-motion and the visible subject's spatial trajectory. This camera-trajectory decomposition naturally extracts synchronized multi-agent motion data within a shared 3D space, completely bypassing the need for multi-camera setups. Second, for precise visual control, we develop the Subject-Aware World Generator to enable appearance-driven simulation conditioned on per-agent identity images. Finally, to ensure both views are grounded in the identical physical reality, we propose World-State Alignment, a per-frame inter-branch cross-attention mechanism inserted at every transformer layer of the video DiT. By jointly synchronizing the denoising process, WSA enforces both static geometric consistency and dynamic motion consistency, encouraging that the shared 3D environment and physical events remain well-aligned across both egocentric views. Extensive experiments demonstrate that MetaWorld achieves superior cross-view consistency and identity fidelity, establishing a highly scalable, physics-driven paradigm for multi-agent video world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。