arXiv:2606.17536cs.CVcs.AI2026-06

用大模型协调多智能体,让驾驶视频生成更一致、更真实。

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

论文配图:OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation
图 1 · 摘自论文原文
  • 大模型指挥三个智能体协同生成带空间信息的视频帧序列。
  • 在nuScenes上实现21.6的BEV mAP和45.7的FVD,优于现有方法。
  • 适合自动驾驶仿真、数据增强及多视角一致性研究者。

面向自动驾驶的生成式世界模型面临两大难题:控制指令异构(自然语言、高精地图、轨迹、相机位姿分属不同表征空间)与事后跨视图融合(单摄像头潜在表示无法捕捉全局三维几何)。我们发现二者根源在于缺乏统一符号化中间语言,在潜在令牌层对齐语言、几何与像素。为此提出DRIVE-CHOREO,一个由大模型协调的多智能体世界模型,将可控多视角视频生成重构为潜在空间中的协同编排。三个Qwen2.5-VL智能体——导演解析用户意图生成结构化WorldScript,制图师将其锚定为空间位置令牌,审计员回传跨视图批评作为辅助监督——共同撰写单一位置感知令牌序列。该序列与多视角视频通过视图-时间排列联合压缩,强制在3D VAE的卷积感受野内保持相机间几何关系。在nuScenes上,DRIVE-CHOREO实现21.6的BEV mAP与45.7的FVD,刷新多视角一致性与检测性能记录;基于其合成数据训练的探测器在真实验证集上提升2.4 NDS,验证了下游实用性。

原文摘要 · Abstract (English)

Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-maps, trajectories, and camera poses reside in incompatible representational spaces, and post-hoc cross-view fusion, where per-camera latents fail to encode global 3-D geometry. We trace both to a single root cause: the absence of a shared symbolic interlingua aligning language, geometry, and pixels at the latent-token level. We present DRIVE-CHOREO, an LLM-choreographed multi-agent world model that recasts controllable multi-view video generation as latent choreography. Three Qwen2.5-VL agents - a Director parsing user intent into a structured WorldScript, a Cartographer grounding it into spatially-anchored layout tokens, and an Auditor feeding cross-view critiques back as auxiliary supervision - jointly author a single position-aware token sequence. This sequence is co-compressed with the multi-view video via a view-time permutation that enforces inter-camera geometry within the convolutional receptive field of a 3-D VAE. On nuScenes, DRIVE-CHOREO sets new state-of-the-art multi-view consistency and BEV mAP (21.6) with competitive FVD (45.7); a detector trained purely on our synthetic data gains +2.4 NDS on the real validation split, validating downstream utility.

自动驾驶多视角生成大模型协同世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。