arXiv:2508.10935cs.CVcs.LG2025-08

用扩散模型提升3D检测伪标签质量,让新类别识别更准。

HQ-OV3D: A High Box Quality Open-World 3D Detection Framework based on Diffision Model

  • 通过跨模态几何一致性生成高质量初始3D框
  • 利用标注类几何先验,用扩散机制逐步优化框精度
  • 在新类别上提升7.37%的mAP,适合自动驾驶等开放场景

传统闭集3D检测框架难以满足自动驾驶等开放世界应用需求。现有开放词汇3D检测方法多采用两阶段流程:先生成伪标签,再进行语义对齐。尽管视觉语言模型(VLMs)显著提升了伪标签语义准确性,但其几何质量(特别是边界框精度)常被忽视。为此,我们提出高框质量开放词汇3D检测框架(HQ-OV3D),专注于生成与精炼高质量伪标签。该框架包含两个核心组件:基于跨模态几何一致性生成初始3D提案的内部模态交叉验证(IMCV)生成器,以及利用标注类别几何先验,通过基于DDIM的去噪机制逐步优化提案的标注类辅助(ACA)去噪器。相较于最先进方法,使用本框架生成的伪标签训练模型,在新类别上实现7.37%的mAP提升,证明了其伪标签的优越性。HQ-OV3D不仅能作为强性能的独立开放词汇3D检测器,还可作为插件式高质量伪标签生成器,集成至现有开放词汇检测或标注流程中。

原文摘要 · Abstract (English)

Traditional closed-set 3D detection frameworks fail to meet the demands of open-world applications like autonomous driving. Existing open-vocabulary 3D detection methods typically adopt a two-stage pipeline consisting of pseudo-label generation followed by semantic alignment. While vision-language models (VLMs) recently have dramatically improved the semantic accuracy of pseudo-labels, their geometric quality, particularly bounding box precision, remains commonly neglected. To address this issue, we propose a High Box Quality Open-Vocabulary 3D Detection (HQ-OV3D) framework, dedicated to generate and refine high-quality pseudo-labels for open-vocabulary classes. The framework comprises two key components: an Intra-Modality Cross-Validated (IMCV) Proposal Generator that utilizes cross-modality geometric consistency to generate high-quality initial 3D proposals, and an Annotated-Class Assisted (ACA) Denoiser that progressively refines 3D proposals by leveraging geometric priors from annotated categories through a DDIM-based denoising mechanism. Compared to the state-of-the-art method, training with pseudo-labels generated by our approach achieves a 7.37% improvement in mAP on novel classes, demonstrating the superior quality of the pseudo-labels produced by our framework. HQ-OV3D can serve not only as a strong standalone open-vocabulary 3D detector but also as a plug-in high-quality pseudo-label generator for existing open-vocabulary detection or annotation pipelines.

3D检测扩散模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。