arXiv:2608.19036cs.CV2026-08

统一建模3D高斯与框,提升自动驾驶场景理解精度

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

论文配图:USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes
图 1 · 摘自论文原文
  • 将3D高斯与边界框作为对齐的潜在序列,联合去噪生成
  • 在nuScenes和VKitti上同时刷新动态重建与3D检测性能
  • 适合做端到端自动驾驶感知与几何建模的研究者

自动驾驶中的空间表征学习旨在将原始视觉信号映射为结构化3D场景表示,其中以对象为中心的边界框和渲染导向的3D高斯原语构成两种互补但独立的层次。现有方法通常将动态重建与实例级感知视为独立任务,尽管其目标均为估计底层3D世界状态。这导致动态重建约束不足,而3D检测缺乏几何依据。为此,我们提出USR-Drive:一种统一的条件生成框架,在仅输入带姿态的多视角驾驶视频时,联合恢复密集动态几何与实例级物体布局。具体地,将密集高斯原语与稀疏3D边界框表示为两个对齐的潜在标记流,并通过统一的多模态扩散Transformer联合去噪。不同于以往将框作为外部条件或用独立模块预测的方法,USR-Drive将其视为相互约束的状态变量,采用统一位置编码(UPE)在共享时空坐标系中对齐异构标记。该统一表征与生成框架使两者相互增强:几何提供密集度量证据支持框预测,框则提供实例级结构先验,有助于保持空间一致性并减少序列3D几何表征的歧义。本方法在nuScenes和VKitti数据集上均实现了动态重建与3D检测的最先进性能。

原文摘要 · Abstract (English)

Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typically treat dynamic reconstruction and instance-level perception as separate tasks, despite their shared goal of estimating the underlying 3D world state. As a result, dynamic reconstruction is under-constrained while 3D detection lacks geometric grounding. To address this gap, we propose USR-Drive, a unified conditional generative framework that, given only posed multi-view driving videos, jointly recovers dense dynamic geometry and instance-level object layouts within a shared scene representation. Specifically, USR-Drive represents dense Gaussian primitives and sparse 3D bounding boxes as two aligned latent token streams and jointly denoises them with a unified multi-modal diffusion Transformer. Unlike prior paradigms that use boxes as external conditions or predict them with detached modules, USR-Drive treats them as mutually constrained state variables with a Unified Positional Encoding (UPE) that aligns heterogeneous tokens within a shared metric spatiotemporal coordinate. Via such unified representation and generative framework, the two modalities reinforce each other: geometry supplies dense metric evidence for box prediction, while boxes provide instance-level structural priors that help preserve spatial consistency and reduce ambiguity in sequential 3D geometric representation. Our approach successfully delivers state-of-the-art results for both dynamic reconstruction and 3D detection on the nuScenes and VKitti datasets.

自动驾驶3D重建扩散模型统一表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。