arXiv:2505.23597cs.CV2025-05CVPR被引 3

用新型卷积层提升树冠语义分割精度,兼顾细节与全局信息。

Bridging Classical and Modern Computer Vision: PerceptiveNet for Tree Crown Semantic Segmentation

  • 引入可训练的对数伽柏卷积层,增强对复杂纹理和尺度变化的感知能力。
  • 在树冠数据集上超越现有模型,跨域测试中仍保持优异表现。
  • 适合遥感图像分析、森林监测等领域的研究者使用。

遥感数据中树冠的精确语义分割对森林管理、生物多样性研究及碳汇量化至关重要。然而,林冠阴影、复杂背景、尺度变化及树种间细微光谱差异使精准分割仍具挑战。相比传统方法,深度学习模型虽能提取判别性特征,却难以捕捉上述复杂性。为此,我们提出PerceptiveNet,结合可训练参数的对数伽柏卷积层与具备宽感受野的主干网络,以提取显著特征并保留全局空间信息。通过大量实验对比对数伽柏、伽柏与标准卷积层的效果,并开展消融实验评估各组件贡献,同时将PerceptiveNet作为新混合CNN-Transformer模型的主干进行验证。结果表明,该模型在树冠数据集上优于当前最优模型,并在两个不同复杂度的航拍场景语义分割基准数据集上实现良好泛化性能。

原文摘要 · Abstract (English)

The accurate semantic segmentation of tree crowns within remotely sensed data is crucial for scientific endeavours such as forest management, biodiversity studies, and carbon sequestration quantification. However, precise segmentation remains challenging due to complexities in the forest canopy, including shadows, intricate backgrounds, scale variations, and subtle spectral differences among tree species. Compared to the traditional methods, Deep Learning models improve accuracy by extracting informative and discriminative features, but often fall short in capturing the aforementioned complexities. To address these challenges, we propose PerceptiveNet, a novel model incorporating a Logarithmic Gabor-parameterised convolutional layer with trainable filter parameters, alongside a backbone that extracts salient features while capturing extensive context and spatial information through a wider receptive field. We investigate the impact of Log-Gabor, Gabor, and standard convolutional layers on semantic segmentation performance through extensive experimentation. Additionally, we conduct an ablation study to assess the contributions of individual layers and their combinations to overall model performance, and we evaluate PerceptiveNet as a backbone within a novel hybrid CNN-Transformer model. Our results outperform state-of-the-art models, demonstrating significant performance improvements on a tree crown dataset while generalising across domains, including two benchmark aerial scene semantic segmentation datasets with varying complexities.

语义分割遥感图像树冠识别深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。