arXiv:2507.04084cs.GRcs.CV2025-07被引 1

通过注意力引导的多尺度局部重建,提升点云自监督学习效果

Attention-Guided Multi-Scale Local Reconstruction for Point Clouds via Masked Autoencoder Self-Supervised Learning

  • 分层重建:低层恢复细节特征,高层处理整体结构
  • 引入局部注意力模块,增强语义特征表达能力
  • 在多个真实数据集上表现优异,适合实际应用

自监督学习已成为点云处理的重要方向。现有模型多聚焦于编码器高层的重构任务,往往忽视了低层局部特征的有效利用,这些特征通常仅用于激活计算而未直接参与重构。为克服此局限,我们提出PointAMaLR,一种新颖的自监督学习框架,通过注意力引导的多尺度局部重构,提升特征表示与处理精度。PointAMaLR在多个局部区域实现分层重构,低层专注于细粒度特征恢复,高层负责粗粒度特征重构,从而支持复杂块间交互。此外,在嵌入层引入局部注意力(LA)模块,以增强语义理解能力。在ModelNet和ShapeNet等基准数据集上的实验表明,PointAMaLR在分类与重构任务中均取得更优性能。在真实场景数据集ScanObjectNN和3D大场景分割数据集S3DIS上,也展现出高度竞争力的表现。结果不仅验证了其在多尺度语义理解中的有效性,也凸显了其在真实场景中的实用性。

原文摘要 · Abstract (English)

Self-supervised learning has emerged as a prominent research direction in point cloud processing. While existing models predominantly concentrate on reconstruction tasks at higher encoder layers, they often neglect the effective utilization of low-level local features, which are typically employed solely for activation computations rather than directly contributing to reconstruction tasks. To overcome this limitation, we introduce PointAMaLR, a novel self-supervised learning framework that enhances feature representation and processing accuracy through attention-guided multi-scale local reconstruction. PointAMaLR implements hierarchical reconstruction across multiple local regions, with lower layers focusing on fine-scale feature restoration while upper layers address coarse-scale feature reconstruction, thereby enabling complex inter-patch interactions. Furthermore, to augment feature representation capabilities, we incorporate a Local Attention (LA) module in the embedding layer to enhance semantic feature understanding. Comprehensive experiments on benchmark datasets ModelNet and ShapeNet demonstrate PointAMaLR's superior accuracy and quality in both classification and reconstruction tasks. Moreover, when evaluated on the real-world dataset ScanObjectNN and the 3D large scene segmentation dataset S3DIS, our model achieves highly competitive performance metrics. These results not only validate PointAMaLR's effectiveness in multi-scale semantic understanding but also underscore its practical applicability in real-world scenarios.

点云处理自监督学习多尺度重建注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。