arXiv:2501.14004cs.CVcs.AI2025-01被引 4

用多任务注意力模型提升城市3D变化检测精度

ME-CPT: Multi-Task Enhanced Cross-Temporal Point Transformer for Urban 3D Change Detection

  • 设计跨时序点云注意力网络,联合提取语义与变化特征
  • 多任务训练缓解类别不平衡,提升变化类型区分度
  • 发布22.5平方公里真实场景数据集,推动领域发展

机载激光雷达(ALS)系统获取的点云提供城市地物精确的三维信息。利用多时相ALS点云可捕捉城市区域的语义变化,在城市规划、应急管理和基础设施维护中具有重要应用价值。现有3D变化检测方法难以高效提取多类语义信息和变化特征,面临三大挑战:(1)跨时相点云空间关系建模困难,影响变化特征提取;(2)变化样本类别不平衡,削弱语义特征区分能力;(3)缺乏真实世界3D语义变化检测数据集。为此,本文提出多任务增强的跨时序点变换器(ME-CPT)。ME-CPT建立不同时相点云间的时空对应关系,通过注意力机制联合提取语义变化特征,促进信息交互与变化比对。同时引入语义分割任务,借助多任务训练策略进一步增强语义特征可区分性,缓解变化类型类别不平衡问题。此外,我们发布了22.5 $km^2$ 的3D语义变化检测数据集,涵盖多样场景,支持全面评估。在多个数据集上的实验表明,所提方法优于现有最先进方法。代码与数据集将在论文接受后开源于https://github.com/zhangluqi0209/ME-CPT。

原文摘要 · Abstract (English)

The point clouds collected by the Airborne Laser Scanning (ALS) system provide accurate 3D information of urban land covers. By utilizing multi-temporal ALS point clouds, semantic changes in urban area can be captured, demonstrating significant potential in urban planning, emergency management, and infrastructure maintenance. Existing 3D change detection methods struggle to efficiently extract multi-class semantic information and change features, still facing the following challenges: (1) the difficulty of accurately modeling cross-temporal point clouds spatial relationships for effective change feature extraction; (2) class imbalance of change samples which hinders distinguishability of semantic features; (3) the lack of real-world datasets for 3D semantic change detection. To resolve these challenges, we propose the Multi-task Enhanced Cross-temporal Point Transformer (ME-CPT) network. ME-CPT establishes spatiotemporal correspondences between point cloud across different epochs and employs attention mechanisms to jointly extract semantic change features, facilitating information exchange and change comparison. Additionally, we incorporate a semantic segmentation task and through the multi-task training strategy, further enhance the distinguishability of semantic features, reducing the impact of class imbalance in change types. Moreover, we release a 22.5 $km^2$ 3D semantic change detection dataset, offering diverse scenes for comprehensive evaluation. Experiments on multiple datasets show that the proposed MT-CPT achieves superior performance compared to existing state-of-the-art methods. The source code and dataset will be released upon acceptance at https://github.com/zhangluqi0209/ME-CPT.

3D变化检测点云处理多任务学习城市建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。