用状态空间模型提升激光雷达3D目标跟踪的效率与精度
MambaTrack3D: A State Space Model Framework for LiDAR-Based Object Tracking under High Temporal Variation
- 基于Mamba设计跨帧传播模块,实现近线性复杂度
- 在KITTI-HTV和nuScenes-HTV上成功率最高提升6.5%
- 适合高时变环境下的实时3D目标跟踪应用
动态室外环境中高时间变化(HTV)给激光雷达点云中的3D单目标跟踪带来挑战。现有基于记忆的跟踪器常面临二次计算复杂度、时间冗余以及几何先验利用不足的问题。为此,我们提出MambaTrack3D,一种面向HTV的新型跟踪框架,基于状态空间模型Mamba构建。具体地,设计了基于Mamba的跨帧传播(MIP)模块,取代传统单帧特征提取,实现近线性复杂度并显式建模历史帧间的空间关系。此外,引入分组特征增强模块(GFEM),在通道级别分离前景与背景语义,有效缓解记忆库中的时间冗余。在KITTI-HTV和nuScenes-HTV基准上的大量实验表明,MambaTrack3D持续优于各类HTV导向及常规场景跟踪器,在中等时间间隔下较HVTrack成功率提升最高达6.5%,精度提升9.5%。在标准KITTI数据集上,其性能仍与顶尖常规场景跟踪器相当,验证了良好的泛化能力。总体而言,MambaTrack3D实现了优异的精度-效率权衡,在专用HTV与常规跟踪场景中均表现稳健。
原文摘要 · Abstract (English)
Dynamic outdoor environments with high temporal variation (HTV) pose significant challenges for 3D single object tracking in LiDAR point clouds. Existing memory-based trackers often suffer from quadratic computational complexity, temporal redundancy, and insufficient exploitation of geometric priors. To address these issues, we propose MambaTrack3D, a novel HTV-oriented tracking framework built upon the state space model Mamba. Specifically, we design a Mamba-based Inter-frame Propagation (MIP) module that replaces conventional single-frame feature extraction with efficient inter-frame propagation, achieving near-linear complexity while explicitly modeling spatial relations across historical frames. Furthermore, a Grouped Feature Enhancement Module (GFEM) is introduced to separate foreground and background semantics at the channel level, thereby mitigating temporal redundancy in the memory bank. Extensive experiments on KITTI-HTV and nuScenes-HTV benchmarks demonstrate that MambaTrack3D consistently outperforms both HTV-oriented and normal-scenario trackers, achieving improvements of up to 6.5 success and 9.5 precision over HVTrack under moderate temporal gaps. On the standard KITTI dataset, MambaTrack3D remains highly competitive with state-of-the-art normal-scenario trackers, confirming its strong generalization ability. Overall, MambaTrack3D achieves a superior accuracy-efficiency trade-off, delivering robust performance across both specialized HTV and conventional tracking scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。