arXiv:2608.16234cs.CV2026-08

用3D高斯表示统一驾驶场景理解与可控生成,支持语言指令编辑。

GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation

论文配图:GaussianDWM++: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation
图 1 · 摘自论文原文
  • 将视觉语言特征直接转为3D高斯基元,构建开放词汇语义场。
  • 通过几何感知适配器实现文本条件下的紧凑世界令牌聚合。
  • 支持天气、车辆等动态元素的指令可控4D编辑,性能领先基准。

驾驶世界模型(DWMs)近年来在生成模型推动下快速发展,但多数方法仅关注条件化场景生成,缺乏显式的3D场景理解、语言接地推理及可控4D编辑能力。现有点云、体素或前视图(BEV)表示难以实现文本信息与3D场景结构的精细对齐。为此,我们提出一个基础特征高斯驱动世界模型,统一场景理解、语言接地推理、可控4D编辑与多模态生成。具体地,设计了一种基础特征高斯分词器,直接将Qwen/SigLIP视觉-语言特征提炼为3D高斯原语,构建紧凑的开放词汇高斯语义场。进一步引入几何感知高斯适配器,结合重要性感知层级选择与文本条件的Perceiver式交叉注意力,将密集高斯原语聚合为紧凑世界令牌。为提升表示兼容性,提出基于KL散度的高斯-图像分布对齐目标,使高斯世界令牌与基础图像令牌对齐。基于对齐的高斯表示,框架支持指令可控的场景编辑,包括天气条件生成与动态车辆操作。在更广泛的驾驶基准上进行的大量实验表明,该方法在场景理解、视觉定位、规划导向推理和可控4D生成任务中均达到当前最优性能。代码与数据集将公开于GitHub。

原文摘要 · Abstract (English)

Driving World Models (DWMs) have recently advanced rapidly with generative models, yet most existing methods mainly focus on conditional scene generation and lack explicit 3D scene understanding, language-grounded reasoning, and controllable 4D editing capabilities. Moreover, commonly used point cloud, occupancy, or BEV representations make it difficult to achieve fine-grained alignment between textual information and the underlying 3D scene structure. To address these limitations, we propose a foundation-feature Gaussian driving world model that unifies scene understanding, language-grounded reasoning, controllable 4D editing, and multi-modal generation within a single framework. Specifically, we introduce a foundation-feature Gaussian tokenizer that directly distills Qwen/SigLIP visual-language features into 3D Gaussian primitives, building a compact open-vocabulary Gaussian semantic field. We further design a geometry-aware Gaussian adapter that combines importance-aware hierarchical selection with text-conditioned Perceiver-style cross-attention to aggregate dense Gaussian primitives into compact world tokens. To improve representation compatibility, we introduce a KL-based Gaussian--image distribution alignment objective that aligns Gaussian world tokens with foundation image tokens. Based on the aligned Gaussian representation, our framework further supports instruction-controllable scene editing, including weather-conditioned generation and dynamic vehicle manipulation. Extensive experiments on broader driving benchmarks demonstrate that our method achieves state-of-the-art performance across scene understanding, visual grounding, planning-oriented reasoning, and controllable 4D generation tasks. We will release the code and datasets publicly on Github.

3D生成高斯扩散语言控制自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。