用语义引导的专家蒸馏,提升罕见物体在纯摄像头3D检测中的识别能力。
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

- 根据语义相似性路由查询到专用专家,分离易混淆类别。
- 通过CLIP对齐2D语义,增强稀有类别的特征区分度。
- 适合关注自动驾驶中罕见但关键物体检测的研究者。
纯摄像头3D目标检测因其低成本和可扩展性成为自动驾驶的有力替代方案,但现有方法多关注整体性能,忽视真实数据集中的严重长尾分布问题。实践中,儿童、婴儿车、应急车辆等安全关键类别极度稀少,导致模型学习偏倚、性能下降。该挑战因类间模糊性(如视觉相似子类)和类内多样性(外观、尺度、姿态、上下文差异大)而加剧,共同阻碍可靠识别。本文提出SemLT3D,一种语义引导的专家蒸馏框架,通过语义先验丰富稀有类别的表征空间。其包含:(1) 语言引导的专家混合模块,依据语义亲和度将3D查询路由至专用专家,提升模型对尾部分布的专一性与混淆类别的解耦能力;(2) 语义投影蒸馏流程,将3D查询与CLIP启发的2D语义对齐,生成跨多种视觉表现更连贯、更具判别性的特征。尽管针对长尾不平衡设计,其语义结构化学习也增强了对广泛外观变化和极端情况的鲁棒性,为更可靠的纯摄像头3D感知提供了原则性进展。
原文摘要 · Abstract (English)
Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily prioritize overall performance while overlooking the severe long-tail imbalance inherent in real-world datasets. In practice, many rare but safety-critical categories such as children, strollers, or emergency vehicles are heavily underrepresented, leading to biased learning and degraded performance. This challenge is further exacerbated by pronounced inter-class ambiguity (e.g., visually similar subclasses) and substantial intra-class diversity (e.g., objects varying widely in appearance, scale, pose, or context), which together hinder reliable long-tail recognition. In this work, we introduce SemLT3D, a Semantic-Guided Expert Distillation framework designed to enrich the representation space for underrepresented classes through semantic priors. SemLT3D consists of: (1) a language-guided mixture-of-experts module that routes 3D queries to specialized experts according to their semantic affinity, enabling the model to better disentangle confusing classes and specialize on tail distributions; and (2) a semantic projection distillation pipeline that aligns 3D queries with CLIP-informed 2D semantics, producing more coherent and discriminative features across diverse visual manifestations. Although motivated by long-tail imbalance, the semantically structured learning in SemLT3D also improves robustness under broader appearance variations and challenging corner cases, offering a principled step toward more reliable camera-only 3D perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。