arXiv:2607.14005cs.CVcs.RO2026-07

M⁴World实现多视角多模态驾驶世界生成与分钟级稳定输出。

M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

论文配图:M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
图 1 · 摘自论文原文
  • 通过多阶段训练实现四步去噪的因果生成,支持分钟级流式输出。
  • 可精确控制物体空间布局与外观,满足交互式操作需求。
  • 适合需要长时序可控模拟的自动驾驶研发与场景增强任务。

驾驶世界生成已成为可扩展自动驾驶仿真核心能力,但现有方法在对象级可控性与长时序稳定性方面仍受限。我们提出 M⁴World,一种多视角多模态生成式驾驶世界模型,可合成未来环绕视图视频流与同步 LiDAR 扫描数据,支持交互式物体操控和稳定的分钟级流式输出。细粒度物体操控通过灵活的条件输入接口实现,支持对单个物体的空间布局与视觉外观进行显式控制。分钟级流式输出的稳定性则通过多阶段训练框架实现,在仅四步去噪步骤下完成在线因果生成,同时保持长时间推演中的世界动态一致性。基于这些组件,我们引入高效的少片段后训练方法及一系列视觉参考条件生成模型,在保留通用生成能力的同时,支持罕见场景的定制化控制。为评估超越真实感的可控性,我们进一步构建自动化 VLM 基判别流水线,评估场景级条件遵循度、视图级物体可控性与跨视图物体一致性。全面实验表明,M⁴World 持续提供高质量生成、精准可控性与稳定分钟级流式输出。结合下游长尾数据增强与场景编辑任务,充分展现了其在可控、可扩展驾驶仿真中的潜力。

原文摘要 · Abstract (English)

Driving-world generation has emerged as a core capability for scalable autonomous-driving simulation, yet existing methods remain limited in object-level controllability and long-horizon stability. We present M$^\text{4}$World, a Multi-view and Multimodal generative driving world model that synthesizes future surround-view video streams and synchronized LiDAR scans while supporting interactive object Manipulation and stable Minute-long streaming. Fine-grained object manipulation is realized through a flexible conditioning interface that supports explicit control over both the spatial layout and visual appearance of individual objects. Stable minute-long streaming, on the other hand, is achieved through a multi-stage training framework that enables online causal generation in only four denoising steps while maintaining coherent world dynamics throughout extended rollouts. Building on these components, we introduce an efficient few-clip post-training as well as a suite of visual reference-conditioned generation models, preserving general generation ability while allowing rare-case customization for long-tail controllability. To assess controllability beyond realism, we further introduce an automated VLM-based judging pipeline that evaluates scene-level condition adherence, view-wise object controllability, and cross-view object consistency. Comprehensive experiments show that M$^\text{4}$World consistently delivers high generation quality, precise controllability, and stable minute-long streaming. Together with downstream long-tail augmentation and scene editing, these results demonstrate the potential of M$^\text{4}$World for controllable, scalable driving simulation.

驾驶仿真多模态生成可控生成长时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。