仅用激光雷达提升复杂室内动态目标检测的召回率
Robust Dynamic Object Detection in Cluttered Indoor Scenes via Learned Spatiotemporal Cues
- 融合时序占用网格与学习的鸟瞰图动态先验
- 在复杂场景中召回率提升28.67%,F1提升18.50%
- 适合自动驾驶在遮挡和近邻干扰下的动态障碍物检测
在复杂室内环境中实现可靠的动态目标检测仍是自主导航的关键挑战。纯几何激光雷达方法依赖聚类和启发式过滤,在动态物体紧邻静态结构或部分可见时易遗漏。视觉增强方法虽提供语义线索,但受限于封闭集检测器和相机视场,对新障碍物和视场外事件鲁棒性差。本文提出一种仅使用激光雷达的框架,将基于时序占用网格的运动分割与学习的鸟瞰图动态先验相融合。融合模块在3D检测可用时优先采用,否则利用学习的动态网格恢复因近邻导致的漏检。实验基于动作捕捉真值数据表明,该方法在高度杂乱环境中相比最先进方法召回率提升28.67%,F1分数提升18.50%,同时保持相近精度与定位误差。
原文摘要 · Abstract (English)
Reliable dynamic object detection in cluttered environments remains a critical challenge for autonomous navigation. Purely geometric LiDAR pipelines that rely on clustering and heuristic filtering can miss dynamic obstacles when they move in close proximity to static structure or are only partially observed. Vision-augmented approaches can provide additional semantic cues, but are often limited by closed-set detectors and camera field-of-view constraints, reducing robustness to novel obstacles and out-of-frustum events. In this work, we present a LiDAR-only framework that fuses temporal occupancy-grid-based motion segmentation with a learned bird's-eye-view (BEV) dynamic prior. A fusion module prioritizes 3D detections when available, while using the learned dynamic grid to recover detections that would otherwise be lost due to proximity-induced false negatives. Experiments with motion-capture ground truth show our method achieves 28.67% higher recall and 18.50% higher F1 score than the state-of-the-art in substantially cluttered environments while maintaining comparable precision and position error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。