针对极稀疏的远距离感知,提出新型融合与注意力机制提升检测精度。
ATN3D: Density-Aware LiDAR-Radar Early 3D Object Detection Under Extreme Sparsity

- 基于密度和雷达信号动态调节多模态融合,减少噪声干扰。
- 在30米外检测性能提升2.09%(浓雾)至3.33%(晴天)。
- 适合自动驾驶中远距离、低密度场景下的早期目标检测。
3D目标检测是自动驾驶和智能交通系统的核心感知任务。长距离探测因感知信号极度稀疏而困难,但在真实道路场景中却十分常见。尽管计算机视觉中超过30米即视为长距离,但实际仅提供约1-2秒的感知与决策时间。在极端稀疏条件下,两大挑战凸显:一是早期多模态融合易丢失稀疏性信息,并引入空单元或误占单元的噪声,降低远距召回率;二是无上下文的均匀通道监督偏好密集近距样本,导致远距小目标优化不足,延迟最早检测时机。本文提出面向稀疏场景的LiDAR-Radar框架ATN3D,包含四项创新:(i) 密度感知早期融合,通过跨模态门控根据每个体素密度与雷达证据动态调整融合策略;(ii) 占用门控邻域聚合,采用环形核仅聚合可信体素;(iii) 证据条件化通道自注意力,根据天气/距离动态调整通道权重;(iv) 距离感知损失,按距离重新平衡分类与定位任务,使训练与分层评估对齐。在VoD基准上,晴天条件下mAP提升3.55%,浓雾下达8.41%;对于>30米物体,晴天增益3.33%,浓雾下2.09%。结果表明,该方法显著提升了道路场景下稀疏感知中的早期可靠长距离检测能力。
原文摘要 · Abstract (English)
3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications. Long-range detection is challenging because sensing evidence is sparse; yet this ``long-range'' scenario is routine in traffic. Although >30m is often labeled long-range in computer vision, on roadways it affords only approx. 1-2s for perception and decision-making. Under such extreme sparsity, two core challenges arise. First, early multimodal fusion tends to discard sparsity information and inject noise from empty or falsely occupied cells, degrading long-range recall. Second, context-agnostic uniform channel supervision favors dense and near-range samples, leaving far and small objects under-optimized, delaying the earliest detection of distant objects. We propose ``Ask The Neighbor'' (ATN3D), a LiDAR-Radar framework tailored for sparse-range conditions. ATN3D introduces (i) Density-aware early fusion with cross-modal gating that conditions fusion on per-voxel density/sparsity and Radar evidence, (ii) Occupancy-gated neighborhood aggregation with circular kernels to aggregate only from credible cells, (iii) Evidence-conditioned channel self-attention to adapt channel weights with weather/range, and (iv) a Range-aware loss that re-balances classification and localization by distance, aligning training with distance-stratified evaluation. On the VoD benchmark across clear and foggy conditions, ATN3D surpasses strong baselines: +3.55% mAP in clear weather and +8.41% mAP under simulated heavy fog; for >30m objects, gains are +3.33% (clear) and +2.09% (heavy fog). These results indicate earlier and more reliable long-range detections under sparse sensing in on-road traffic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。