统一建模3D高斯与框,提升自动驾驶场景理解精度
USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

- 将3D高斯与边界框作为对齐的潜在序列,联合去噪生成
- 在nuScenes和VKitti上同时刷新动态重建与3D检测性能
- 适合做端到端自动驾驶感知与几何建模的研究者
自动驾驶中的空间表征学习旨在将原始视觉信号映射为结构化3D场景表示,其中以对象为中心的边界框和渲染导向的3D高斯原语构成两种互补但独立的层次。现有方法通常将动态重建与实例级感知视为独立任务,尽管其目标均为估计底层3D世界状态。这导致动态重建约束不足,而3D检测缺乏几何依据。为此,我们提出USR-Drive:一种统一的条件生成框架,在仅输入带姿态的多视角驾驶视频时,联合恢复密集动态几何与实例级物体布局。具体地,将密集高斯原语与稀疏3D边界框表示为两个对齐的潜在标记流,并通过统一的多模态扩散Transformer联合去噪。不同于以往将框作为外部条件或用独立模块预测的方法,USR-Drive将其视为相互约束的状态变量,采用统一位置编码(UPE)在共享时空坐标系中对齐异构标记。该统一表征与生成框架使两者相互增强:几何提供密集度量证据支持框预测,框则提供实例级结构先验,有助于保持空间一致性并减少序列3D几何表征的歧义。本方法在nuScenes和VKitti数据集上均实现了动态重建与3D检测的最先进性能。
原文摘要 · Abstract (English)
Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typically treat dynamic reconstruction and instance-level perception as separate tasks, despite their shared goal of estimating the underlying 3D world state. As a result, dynamic reconstruction is under-constrained while 3D detection lacks geometric grounding. To address this gap, we propose USR-Drive, a unified conditional generative framework that, given only posed multi-view driving videos, jointly recovers dense dynamic geometry and instance-level object layouts within a shared scene representation. Specifically, USR-Drive represents dense Gaussian primitives and sparse 3D bounding boxes as two aligned latent token streams and jointly denoises them with a unified multi-modal diffusion Transformer. Unlike prior paradigms that use boxes as external conditions or predict them with detached modules, USR-Drive treats them as mutually constrained state variables with a Unified Positional Encoding (UPE) that aligns heterogeneous tokens within a shared metric spatiotemporal coordinate. Via such unified representation and generative framework, the two modalities reinforce each other: geometry supplies dense metric evidence for box prediction, while boxes provide instance-level structural priors that help preserve spatial consistency and reduce ambiguity in sequential 3D geometric representation. Our approach successfully delivers state-of-the-art results for both dynamic reconstruction and 3D detection on the nuScenes and VKitti datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。