arXiv:2604.02903cs.CVcs.AI2026-04

用射线对齐序列化提升远距离3D目标检测的上下文建模能力

RayMamba: Ray-Aligned Serialization for Long-Range 3D Object Detection

  • 按激光射线方向组织稀疏体素,保持空间连续性
  • 在nuScenes上远距检测最高提升2.49点mAP
  • 兼容单模态与多模态检测器,轻量易部署

远距离3D目标检测因激光雷达观测在远场变得极度稀疏和碎片化,导致现有检测器难以可靠建模上下文。尽管基于状态空间模型(SSM)的方法提升了长程建模效率,但其效果仍受限于通用序列化策略,无法保留稀疏场景中的有意义上下文邻域。为此,我们提出RayMamba,一种面向体素化3D检测器的几何感知即插即用增强模块。RayMamba通过射线对齐序列化策略,将稀疏体素按扇区有序排列,保留方向连续性和遮挡相关上下文,便于后续Mamba模型建模。该方法兼容单雷达与多模态检测器,仅引入轻微计算开销。在nuScenes和Argoverse 2上的大量实验表明,其在多个强基线中均带来一致性能提升。尤其在nuScenes的40–50米远距区间,最大提升达2.49 mAP和1.59 NDS;在Argoverse 2上,VoxelNeXt的mAP从30.3提升至31.2。

原文摘要 · Abstract (English)

Long-range 3D object detection remains challenging because LiDAR observations become highly sparse and fragmented in the far field, making reliable context modeling difficult for existing detectors. To address this issue, recent state space model (SSM)-based methods have improved long-range modeling efficiency. However, their effectiveness is still limited by generic serialization strategies that fail to preserve meaningful contextual neighborhoods in sparse scenes. To address this issue, we propose RayMamba, a geometry-aware plug-and-play enhancement for voxel-based 3D detectors. RayMamba organizes sparse voxels into sector-wise ordered sequences through a ray-aligned serialization strategy, which preserves directional continuity and occlusion-related context for subsequent Mamba-based modeling. It is compatible with both LiDAR-only and multimodal detectors, while introducing only modest overhead. Extensive experiments on nuScenes and Argoverse 2 demonstrate consistent improvements across strong baselines. In particular, RayMamba achieves up to 2.49 mAP and 1.59 NDS gain in the challenging 40--50 m range on nuScenes, and further improves VoxelNeXt on Argoverse 2 from 30.3 to 31.2 mAP.

3D检测激光雷达序列建模Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。