arXiv:2508.03252cs.CV2025-08AAAI被引 12

用可拆卸扩散框架实现高效鲁棒的3D目标检测

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

  • 在潜在空间用轻量去噪网络学习去噪过程
  • 单步推理下仍达领先性能,大幅提高效率
  • 适合需要实时性与抗干扰能力的自动驾驶场景

去噪扩散概率模型(DDPM)在鲁棒3D目标检测中表现优异。现有方法依赖3D框的score matching或预训练扩散先验,通常需多步迭代推理,影响效率。为此,我们提出一种基于可拆卸潜在框架(DLF)的单阶段全稀疏3D目标检测网络RSDNet。RSDNet通过多级去噪自编码器(DAEs)在潜在特征空间学习去噪过程,有效理解多层级扰动下的场景分布,实现鲁棒可靠检测。同时,重构了DDPM的加噪与去噪机制,使DLF能生成多类型、多层级噪声样本与目标,增强对多种扰动的鲁棒性。引入语义-几何条件引导,感知物体边界与形状,缓解稀疏表示中的中心特征缺失问题,支持全稀疏检测流程。此外,DLF的可拆卸去噪网络设计使RSDNet在推理时仅需单步完成检测,显著提升效率。大量公开基准测试表明,RSDNet性能超越现有方法,达到当前最优水平。

原文摘要 · Abstract (English)

Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a Robust single-stage fully Sparse 3D object Detection Network with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection.

3D检测扩散模型稀疏检测自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。