让多个智能体在统一世界中生成连贯视频,支持交互与视角一致性。
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

- 通过四视角拼接与跨智能体注意力,实现多智能体共享世界建模。
- 支持49帧长视频生成,动态智能体位置感知准确,重叠区域一致性强。
- 基于CARLA构建大规模交互数据集,适用于自动驾驶等场景模拟。
本文提出ShareVerse,一个支持多智能体共享世界建模的视频生成框架,填补了现有方法在统一共享世界构建与多智能体交互支持方面的空白。该框架利用大视频模型的生成能力,集成三项关键创新:1)在CARLA仿真平台构建大规模多智能体交互世界数据集,包含多样场景、天气条件及交互轨迹,并配有每智能体前后左右四视角视频与相机数据;2)提出四视角视频空间拼接策略,用于建模更广环境并保证多视角几何一致性;3)将跨智能体注意力模块融入预训练视频模型,实现时空信息在智能体间的交互传递,确保重叠区域共享世界的一致性及非重叠区域的合理生成。ShareVerse支持49帧大规模视频生成,能准确感知动态智能体位置,实现一致的共享世界建模。
原文摘要 · Abstract (English)
This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent interaction. ShareVerse leverages the generation capability of large video models and integrates three key innovations: 1) A dataset for large-scale multi-agent interactive world modeling is built on the CARLA simulation platform, featuring diverse scenes, weather conditions, and interactive trajectories with paired multi-view videos (front/ rear/ left/ right views per agent) and camera data. 2) We propose a spatial concatenation strategy for four-view videos of independent agents to model a broader environment and to ensure internal multi-view geometric consistency. 3) We integrate cross-agent attention blocks into the pretrained video model, which enable interactive transmission of spatial-temporal information across agents, guaranteeing shared world consistency in overlapping regions and reasonable generation in non-overlapping regions. ShareVerse, which supports 49-frame large-scale video generation, accurately perceives the position of dynamic agents and achieves consistent shared world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。