对比三种传感器在室内外环境下检测人的表现,融合方案最稳
A Comparative Study of 3D Person Detection: Sensor Modalities and Robustness in Diverse Indoor and Outdoor Environments

- 用相机、激光雷达及融合模型做3D人检测对比
- 融合方法在遮挡和远距离下准确率最高,提升12.3%
- 适合机器人、监控等需要高鲁棒性的场景
精准的3D人体检测对机器人、工业监控和安防等应用的安全至关重要。本文系统评估了仅用相机、仅用激光雷达以及相机-激光雷达融合三种方式在多样室内与室外场景中的3D人体检测性能。基于JRDB数据集,对比BEVDepth(相机)、PointPillars(激光雷达)和DAL(融合)三类代表性模型,并分析其在不同遮挡程度与距离下的表现。结果表明,融合方法在复杂场景中持续优于单模态模型。进一步研究传感器损坏与错位情况发现,虽然DAL具备更强抗干扰能力,但仍对激光雷达错位和特定噪声敏感;而相机模型BEVDepth性能最低,受遮挡、距离和噪声影响最大。研究强调传感器融合对提升3D人体检测的重要性,也指出当前系统仍存在固有脆弱性,需持续改进。
原文摘要 · Abstract (English)
Accurate 3D person detection is critical for safety in applications such as robotics, industrial monitoring, and surveillance. This work presents a systematic evaluation of 3D person detection using camera-only, LiDAR-only, and camera-LiDAR fusion. While most existing research focuses on autonomous driving, we explore detection performance and robustness in diverse indoor and outdoor scenes using the JRDB dataset. We compare three representative models - BEVDepth (camera), PointPillars (LiDAR), and DAL (camera-LiDAR fusion) - and analyze their behavior under varying occlusion and distance levels. Our results show that the fusion-based approach consistently outperforms single-modality models, particularly in challenging scenarios. We further investigate robustness against sensor corruptions and misalignments, revealing that while DAL offers improved resilience, it remains sensitive to sensor misalignment and certain LiDAR-based corruptions. In contrast, the camera-based BEVDepth model showed the lowest performance and was most affected by occlusion, distance, and noise. Our findings highlight the importance of utilizing sensor fusion for enhanced 3D person detection, while also underscoring the need for ongoing research to address the vulnerabilities inherent in these systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。