用双曲几何提升多模态3D目标检测的融合效果
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection

- 引入双曲空间几何约束,优化图像与点云特征融合
- 在KITTI和nuScenes上实现更高精度与更低计算开销的平衡
- 适合需要高鲁棒性多模态感知的自动驾驶场景
跨模态知识蒸馏已成为融合点云与图像特征的有效策略。然而,模态异质性、空间错位及多模态表示危机常限制现有方法效率。为此,我们提出一种基于双曲约束的跨模态蒸馏框架HGC-Det,包含图像分支与点云分支。点云分支含三个核心组件:2D语义引导体素优化(SGVO)、双曲几何约束跨模态特征传输(HFT)和基于特征聚合的几何优化(FAGO)。SGVO利用图像分支的语义线索自适应优化3D空间表示,缓解融合不足问题;HFT利用双曲空间的内在几何特性,减轻高维图像特征与低维点云特征融合时的语义损失;FAGO补偿SGVO可能引入的空间特征退化。在SUN RGB-D、ARKitScenes(室内)及KITTI、nuScenes(室外)数据集上的实验表明,本方法在检测精度与计算成本间取得更优权衡。
原文摘要 · Abstract (English)
Cross-modal knowledge distillation has emerged as an effective strategy for integrating point cloud and image features in 3D perception tasks. However, the modality heterogeneity, spatial misalignment, and the representation crisis of multiple modalities often limit the efficient of these cross-modal distillation methods. To address these limitations in existing approaches, we propose a hyperbolic constrained cross-modal distillation method for multimodal 3D object detection (HGC-Det). The proposed HGC-Det framework includes an image branch and a point cloud branch to extract semantic features from two different modalities. The point cloud branch comprises three core components: a 2D semantic-guided voxel optimization component (SGVO), a hyperbolic geometry constrained cross-modal feature transfer component (HFT), and a feature aggregation-based geometry optimization component (FAGO). Specifically, the SGVO component adaptively refines the spatial representation of the 3D branch by leveraging semantic cues from the image branch, thereby mitigating the issue of inadequate representation fusion. The HFT component exploits the intrinsic geometric properties of hyperbolic space to alleviate semantic loss during the fusion of high-dimensional image features and low-dimensional point cloud features. Finally, the FAGO compensates for potential spatial feature degradation introduced by the 2D semantic-guided voxel optimization component. Extensive experiments on indoor datasets (SUN RGB-D, ARKitScenes) and outdoor datasets (KITTI, nuScenes) demonstrate that our method achieves a better trade-off between detection accuracy and computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。