arXiv:2507.05260cs.CVcs.LG2025-07ICCV被引 10

通过跨视角与长时序融合,提升激光雷达表征学习效果

Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations

论文配图:Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR Representations
图 1 · 摘自论文原文
  • 引入跨视角对齐与长时序特征传播机制,增强时空一致性
  • 在多个主流基准上实现显著性能提升,3D目标检测精度提高2.1%
  • 适合自动驾驶领域研究者,尤其关注自监督表征学习的场景

激光雷达表征学习旨在从大规模、易获取的数据中提取丰富的结构与语义信息,降低对昂贵人工标注的依赖。然而现有方法常忽视激光雷达序列中的固有时空线索,限制了其有效性。本文提出LiMA框架,一种新型长时序图像到激光雷达记忆聚合方法,显式捕捉更远距离的时间相关性以增强激光雷达表征学习。该框架包含三个核心组件:1)跨视角聚合模块,对齐并融合相邻相机视图的重叠区域,构建更统一且无冗余的记忆库;2)长时序特征传播机制,高效对齐与整合多帧图像特征,强化激光雷达表征学习中的时间连贯性;3)跨序列记忆对齐策略,强制不同驾驶序列间的一致性,提升对未见环境的泛化能力。LiMA保持高预训练效率,下游任务中不增加计算开销。在主流激光雷达感知基准上的大量实验表明,其显著提升了激光雷达语义分割与3D目标检测性能。代码已公开,供后续研究使用。

原文摘要 · Abstract (English)

LiDAR representation learning aims to extract rich structural and semantic information from large-scale, readily available datasets, reducing reliance on costly human annotations. However, existing LiDAR representation strategies often overlook the inherent spatiotemporal cues in LiDAR sequences, limiting their effectiveness. In this work, we propose LiMA, a novel long-term image-to-LiDAR Memory Aggregation framework that explicitly captures longer range temporal correlations to enhance LiDAR representation learning. LiMA comprises three key components: 1) a Cross-View Aggregation module that aligns and fuses overlapping regions across neighboring camera views, constructing a more unified and redundancy-free memory bank; 2) a Long-Term Feature Propagation mechanism that efficiently aligns and integrates multi-frame image features, reinforcing temporal coherence during LiDAR representation learning; and 3) a Cross-Sequence Memory Alignment strategy that enforces consistency across driving sequences, improving generalization to unseen environments. LiMA maintains high pretraining efficiency and incurs no additional computational overhead during downstream tasks. Extensive experiments on mainstream LiDAR-based perception benchmarks demonstrate that LiMA significantly improves both LiDAR semantic segmentation and 3D object detection. We hope this work inspires more effective pretraining paradigms for autonomous driving. The code has be made publicly accessible for future research.

激光雷达自监督学习自动驾驶表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。