arXiv:2606.09143cs.CV2026-06

针对无人机视角下物体遮挡问题,提出显式建模遮挡强度的多模态融合方法。

CAMF-Det: Closure-Aware Multimodal Fusion for LiDAR-Camera 3D Object Detection on UAV Platforms

论文配图:CAMF-Det: Closure-Aware Multimodal Fusion for LiDAR-Camera 3D Object Detection on UAV Platforms
图 1 · 摘自论文原文
  • 基于物理模型构建激光雷达与摄像头的双模态遮挡强度真值图
  • 在两个自建无人机数据集上,硬级别检测精度提升超4.8%
  • 适用于复杂遮挡环境下无人机多模态3D目标检测

基于激光雷达与摄像头的多模态3D目标检测在地面车辆场景中表现优异,但在无人机平台尚未深入探索。在无人机俯视场景中,树冠频繁导致地面物体遮挡,造成空间变化且模态依赖的信息退化。现有融合框架未显式建模此类遮挡,也未将遮挡感知融入检测流程,限制了其在遮挡场景下的性能。为此,本文提出CAMF-Det,一种面向无人机平台的遮挡感知多模态融合框架。首先,通过受物理启发的贝叶-兰伯特公式和建筑掩膜修正,离线构建双模态遮挡强度真值图。其次,利用真值图监督,训练网络实现单帧推理时的在线遮挡强度预测。最后,将真值与预测的遮挡强度注入数据增强、特征编码、多模态融合和检测头,实现对空间变化与模态依赖退化的自适应检测。在两个自建无人机多模态数据集SI3D-DI和SI3D-DII上的实验表明,CAMF-Det在所有难度级别均取得最优性能,硬级别mAP$_{\mathrm{BEV}}$分别较最优对比方法提升9.43%和4.88%。结果验证了显式遮挡先验建模与利用的有效性,显著提升了无人机场景下多模态3D检测的鲁棒性。

原文摘要 · Abstract (English)

Multimodal 3D object detection based on LiDAR and cameras has demonstrated excellent performance in ground-vehicle scenarios, but has not been explored for Unmanned Aerial Vehicle (UAV) platforms. In UAV top-down scenes, frequent groundobject occlusion dominated by tree canopies causes spatially varying and modality-dependent information degradation. Existing multimodal fusion frameworks neither explicitly model such ground-object occlusion nor embed occlusion awareness into the detection pipeline, limiting their performance in occluded UAV scenes. To address these challenges, we propose CAMF-Det, a closure-aware multimodal fusion framework for LiDAR-camera 3D object detection on UAV platforms, which derives dual-modal occlusion intensity through physics-inspired modeling and embeds them as priors throughout the detection pipeline. First, a dual-modal closure modeling module explicitly constructs occlusion intensity ground truth for both modalities offline via a Beer-Lambert-inspired formulation and building-mask correction. Second, using these ground-truth maps as supervision, a dual-modal prediction network converts the offline modeling results into online occlusion intensity predictions under single-frame inference. Third, both ground-truth and predicted occlusion intensity are injected into data augmentation, feature encoding, multimodal fusion, and detection head, enabling adaptive detection under spatially varying and modality-dependent information degradation. Experiments on two self-built UAV-based multimodal datasets, SI3D-DI and SI3D-DII, demonstrate that CAMF-Det achieves the best performance across all difficulty levels, with hard-level mAP$_{\mathrm{BEV}}$ improvements of 9.43% and 4.88% over the best competing methods, respectively. These results confirm the effectiveness of explicit occlusion prior modeling and exploitation for robust multimodal 3D detection in UAV scenes.

多模态融合无人机检测遮挡建模3D目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。