在真实航拍条件下评估主流点云模型,揭示空中激光雷达分割的挑战与性能差异。
Benchmarking Deep Learning Models for Aerial LiDAR Point Cloud Semantic Segmentation under Real Acquisition Conditions: A Case Study in Navarre
- 对比4种先进模型在真实航拍数据上的表现,聚焦几何多样性和类别不平衡问题。
- 所有模型总体准确率超93%,KPConv平均交并比达78.51%最优,车辆类识别最弱。
- 适合从事航空遥感、城市建模或点云算法落地的研究者参考。
深度学习在3D语义分割上取得显著进展,但多数模型基于室内或地面数据集,其在真实航拍条件下的表现仍不充分。为填补空白,本研究在西班牙纳瓦拉地区获取的大规模航拍LiDAR数据集上,对四种代表性深度学习模型(KPConv、RandLA-Net、Superpoint Transformer、Point Transformer V3)进行基准测试,覆盖城市、乡村和工业等异质地貌。针对地表、植被、建筑和车辆等五类常见地物进行评估,揭示了空中数据中类别不平衡与几何变异性的固有挑战。结果表明,所有模型整体准确率均超过93%,其中KPConv在各类别上表现稳定,平均交并比达78.51%最高;Point Transformer V3在稀有目标车辆类上表现最佳(IoU 75.11%),而Superpoint Transformer与RandLA-Net则以牺牲分割鲁棒性换取计算效率。
原文摘要 · Abstract (English)
Recent advances in deep learning have significantly improved 3D semantic segmentation, but most models focus on indoor or terrestrial datasets. Their behavior under real aerial acquisition conditions remains insufficiently explored, and although a few studies have addressed similar scenarios, they differ in dataset design, acquisition conditions, and model selection. To address this gap, we conduct an experimental benchmark evaluating several state-of-the-art architectures on a large-scale aerial LiDAR dataset acquired under operational flight conditions in Navarre, Spain, covering heterogeneous urban, rural, and industrial landscapes. This study compares four representative deep learning models, including KPConv, RandLA-Net, Superpoint Transformer, and Point Transformer V3, across five semantic classes commonly found in airborne surveys, such as ground, vegetation, buildings, and vehicles, highlighting the inherent challenges of class imbalance and geometric variability in aerial data. Results show that all tested models achieve high overall accuracy exceeding 93%, with KPConv attaining the highest mean IoU (78.51%) through consistent performance across classes, particularly on challenging and underrepresented categories. Point Transformer V3 demonstrates superior performance on the underrepresented vehicle class (75.11% IoU), while Superpoint Transformer and RandLA-Net trade off segmentation robustness for computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。