arXiv:2501.05095cs.CVcs.AI2025-01被引 2

构建大规模机载激光雷达数据集,提升森林与城市场景的模型性能。

Advancing ALS Applications with Large-Scale Pre-training: Dataset Development and Downstream Assessment

  • 基于地表覆盖和高程图设计地理采样策略,高效构建美国本土激光雷达数据集。
  • 预训练模型在树种分类等任务上显著优于从零训练,验证了迁移学习有效性。
  • 地理采样比随机采样更有效,适合遥感领域大模型预训练应用。

预训练-微调范式已革新卫星遥感应用,但在机载激光扫描(ALS)领域仍研究不足。本文构建覆盖美国本土的大型ALS点云数据集,数据来自美国地质调查局3D高程计划。为兼顾多样性与效率,提出基于地表覆盖图与数字高程模型的地理采样方法。采用当前先进的BEV-MAE自监督模型在该数据集上进行预训练,并用于树种分类、地形场景识别与点云语义分割等下游任务。结果表明,预训练模型在所有任务中均显著优于从零开始训练的模型,证明了所学表示的可迁移性。同时,使用地理采样扩展数据集持续提升性能,而随机采样则无法实现类似增益。这些发现凸显了该数据集与采样策略在预训练范式中的价值。代码与预训练模型将公开于 https://github.com/martianxiu/ALS_pretraining。

原文摘要 · Abstract (English)

The pre-training and fine-tuning paradigm has revolutionized satellite remote sensing applications. However, this approach remains largely underexplored for airborne laser scanning (ALS), an important technology for applications such as forest management and urban planning. In this study, we address this gap by constructing a large-scale ALS point cloud dataset and evaluating its impact on downstream applications. Our dataset comprises ALS point clouds collected across the contiguous United States, provided by the United States Geological Survey's 3D Elevation Program. To ensure efficient data collection while capturing diverse land cover and terrain types, we introduce a geospatial sampling method that selects point cloud tiles based on land cover maps and digital elevation models. As a baseline self-supervised learning model, we adopt BEV-MAE, a state-of-the-art masked autoencoder for 3D outdoor point clouds, and pre-train it on the constructed dataset. The pre-trained models are subsequently fine-tuned for downstream tasks, including tree species classification, terrain scene recognition, and point cloud semantic segmentation. Our results show that the pre-trained models significantly outperform their scratch counterparts across all downstream tasks, demonstrating the transferability of the representations learned from the proposed dataset. Furthermore, we observe that scaling the dataset using our geospatial sampling method consistently enhances performance, whereas pre-training on datasets constructed with random sampling fails to achieve similar improvements. These findings highlight the utility of the constructed dataset and the effectiveness of our sampling strategy in the pre-training and fine-tuning paradigm. The source code and pre-trained models will be made publicly available at \url{https://github.com/martianxiu/ALS_pretraining}.

激光雷达预训练遥感自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。