用少量标注实现3D检测模型跨域迁移,提升真实场景适应能力
From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning
- 结合2D语义与3D空间推理,用视觉语言模型传递可迁移语义
- 在少样本条件下实现对新类别和常见类别的稳定检测,性能优于基线
- 适合需要快速部署到新环境的自动驾驶与机器人场景
基于激光雷达的3D目标检测模型常因现有数据集对象多样性不足,难以泛化至真实场景。为此,我们首次提出3D目标检测中的广义跨域少样本(GCFS)任务,旨在仅用少量标注将源预训练模型适配到新域的常见类与新类别。我们提出统一框架,通过融合2D开放集语义与3D空间推理,在有限监督下学习稳定的目标语义。具体地,图像引导的多模态融合借助视觉语言模型将可迁移的2D语义线索注入3D流程;物理感知的框搜索则利用激光雷达先验增强2D到3D对齐。为从稀疏数据中捕捉类别特异性语义,进一步引入对比增强原型学习,将少样本实例编码为判别性语义锚点,稳定表示学习。在GCFS基准上的大量实验表明,该方法在真实部署场景中具有显著有效性与通用性。
原文摘要 · Abstract (English)
LiDAR-based 3D object detection models often struggle to generalize to real-world environments due to limited object diversity in existing datasets. To tackle it, we introduce the first generalized cross-domain few-shot (GCFS) task in 3D object detection, aiming to adapt a source-pretrained model to both common and novel classes in a new domain with only few-shot annotations. We propose a unified framework that learns stable target semantics under limited supervision by bridging 2D open-set semantics with 3D spatial reasoning. Specifically, an image-guided multi-modal fusion injects transferable 2D semantic cues into the 3D pipeline via vision-language models, while a physically-aware box search enhances 2D-to-3D alignment via LiDAR priors. To capture class-specific semantics from sparse data, we further introduce contrastive-enhanced prototype learning, which encodes few-shot instances into discriminative semantic anchors and stabilizes representation learning. Extensive experiments on GCFS benchmarks demonstrate the effectiveness and generality of our approach in realistic deployment settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。