arXiv:2412.03054cs.CV2024-12NeurIPS被引 4

通过预测点云未来帧,实现无监督3D表征学习

TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR Perception

  • 用时序预测代替对比学习,构建跨帧3D嵌入与时间神经场
  • 在NuScenes等数据集上相比最优无监督方法提升90%
  • 适合做自动驾驶点云感知的预训练,尤其关注运动信息

标注激光雷达点云耗时耗能,促使近期无监督3D表征学习方法通过预训练权重减轻标注负担。现有工作大多聚焦单帧点云,忽视了天然蕴含物体运动与语义的时间序列信息。为此,本文提出TREND(Temporal REndering with Neural fielD),通过无监督方式预测未来观测来学习3D表征。不同于传统的对比学习或掩码自编码范式,TREND采用循环嵌入机制生成跨时间的3D嵌入,并利用时间神经场表示3D场景,通过可微渲染计算损失。据我们所知,TREND是首个基于时序预测的无监督3D表征学习方法。在NuScenes、Once和Waymo等主流数据集的下游3D目标检测任务上评估显示,TREND相比先前最优无监督预训练方法性能提升最高达90%,且在不同下游模型和数据集上普遍表现更优,证明了时序预测对激光雷达感知的有效性。

原文摘要 · Abstract (English)

Labeling LiDAR point clouds is notoriously time-and-energy-consuming, which spurs recent unsupervised 3D representation learning methods to alleviate the labeling burden in LiDAR perception via pretrained weights. Almost all existing work focus on a single frame of LiDAR point cloud and neglect the temporal LiDAR sequence, which naturally accounts for object motion (and their semantics). Instead, we propose TREND, namely Temporal REndering with Neural fielD, to learn 3D representation via forecasting the future observation in an unsupervised manner. Unlike existing work that follows conventional contrastive learning or masked auto encoding paradigms, TREND integrates forecasting for 3D pre-training through a Recurrent Embedding scheme to generate 3D embedding across time and a Temporal Neural Field to represent the 3D scene, through which we compute the loss using differentiable rendering. To our best knowledge, TREND is the first work on temporal forecasting for unsupervised 3D representation learning. We evaluate TREND on downstream 3D object detection tasks on popular datasets, including NuScenes, Once and Waymo. Experiment results show that TREND brings up to 90% more improvement as compared to previous SOTA unsupervised 3D pre-training methods and generally improve different downstream models across datasets, demonstrating that indeed temporal forecasting brings improvement for LiDAR perception.

3D感知无监督学习时序建模激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。