arXiv:2512.14639cs.CV2025-12被引 1

融合CNN与Transformer的模型提升冰川断裂线分割精度

AMD-HookNet++: Evolution of AMD-HookNet with Hybrid CNN-Transformer Feature Enhancement for Glacier Calving Front Segmentation

  • 采用双分支结构:CNN保细节,Transformer捕获长程依赖
  • 在CaFFe数据集上达78.2%交并比,95%豪斯多夫距离1318米
  • 适合关注冰川变化监测与遥感图像分割的研究者

冰川与冰架前端动态显著影响冰盖质量平衡和沿海海平面。为有效监测冰川状态,需持续估计冰川断裂线的位置变化。AMD-HookNet首次提出纯双分支卷积神经网络用于冰川分割。然而,卷积操作的局部性与平移不变性虽利于捕捉低层细节,却限制了模型保持长程依赖的能力。本研究提出AMD-HookNet++,一种新型混合CNN-Transformer特征增强方法,用于合成孔径雷达图像中冰川分割与断裂线描绘。其双分支结构包括基于Transformer的上下文分支以捕捉长程依赖,提供更大视野的全局上下文信息;以及基于CNN的目标分支以保留局部细节。为强化连接的混合特征表示,设计了增强的空间-通道注意力模块,通过动态调整空间与通道维度的特征关系,促进两分支间交互。此外,开发像素级对比深度监督,将像素级度量学习融入冰川分割优化。在具有挑战性的冰川分割基准数据集CaFFe上进行大量实验与全面定量定性分析表明,AMD-HookNet++以78.2%的交并比和1,318米的HD95达到新最佳性能,同时保持367米的竞争力平均距离误差(MDE)。更重要的是,该混合模型生成更平滑的断裂线轮廓,解决了纯Transformer方法常见的锯齿状边缘问题。

原文摘要 · Abstract (English)

The dynamics of glaciers and ice shelf fronts significantly impact the mass balance of ice sheets and coastal sea levels. To effectively monitor glacier conditions, it is crucial to consistently estimate positional shifts of glacier calving fronts. AMD-HookNet firstly introduces a pure two-branch convolutional neural network (CNN) for glacier segmentation. Yet, the local nature and translational invariance of convolution operations, while beneficial for capturing low-level details, restricts the model ability to maintain long-range dependencies. In this study, we propose AMD-HookNet++, a novel advanced hybrid CNN-Transformer feature enhancement method for segmenting glaciers and delineating calving fronts in synthetic aperture radar images. Our hybrid structure consists of two branches: a Transformer-based context branch to capture long-range dependencies, which provides global contextual information in a larger view, and a CNN-based target branch to preserve local details. To strengthen the representation of the connected hybrid features, we devise an enhanced spatial-channel attention module to foster interactions between the hybrid CNN-Transformer branches through dynamically adjusting the token relationships from both spatial and channel perspectives. Additionally, we develop a pixel-to-pixel contrastive deep supervision to optimize our hybrid model by integrating pixelwise metric learning into glacier segmentation. Through extensive experiments and comprehensive quantitative and qualitative analyses on the challenging glacier segmentation benchmark dataset CaFFe, we show that AMD-HookNet++ sets a new state of the art with an IoU of 78.2 and a HD95 of 1,318 m, while maintaining a competitive MDE of 367 m. More importantly, our hybrid model produces smoother delineations of calving fronts, resolving the issue of jagged edges typically seen in pure Transformer-based approaches.

冰川分割混合模型遥感图像目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。