arXiv:2502.15199cs.CV2025-02被引 53

针对遥感城市图像中物体尺度多变问题,提出可学习适配器提升SAM分割精度

UrbanSAM: Learning Invariance-Inspired Adapters for Segment Anything Models in Urban Construction

  • 设计基于多分辨率理论的可学习提示器与缩放适配器,实现尺度不变特征提取
  • 在包含建筑、道路、水体的全球数据集上,对尺度变化物体分割准确率显著提升
  • 适合从事遥感影像分析、城市监测的科研与工程人员使用

从遥感图像中提取和分割目标是城市环境监测中的关键但极具挑战的任务。城市形态固有复杂,物体形状不规则且尺度多样。遥感数据源在传感器、平台和模态上的异质性及尺度差异进一步加剧了精确分割的难度。尽管分割一切模型(SAM)在复杂场景分割中展现出巨大潜力,但其在处理形态多变物体时仍受限于人工交互式提示。为此,我们提出UrbanSAM,一种专为复杂城市环境设计的SAM定制版本,以应对遥感观测中的尺度效应。受多分辨率分析(MRA)理论启发,UrbanSAM引入一种新型可学习提示器,配备遵循不变性准则的Uscaling-Adapter,使模型能捕捉物体的多尺度上下文信息,并理论上适应任意尺度变化。此外,通过掩码交叉注意力操作对齐Uscaling-Adapter与主干编码器的特征,使主干编码器继承适配器的多尺度聚合能力。这种协同机制提升了分割性能,输出更强大、更准确的结果,且由学习到的适配器支撑。大量实验表明,所提UrbanSAM在全球尺度数据集上具备优异的灵活性和分割性能,涵盖建筑、道路、水体等尺度变化的城市场景。

原文摘要 · Abstract (English)

Object extraction and segmentation from remote sensing (RS) images is a critical yet challenging task in urban environment monitoring. Urban morphology is inherently complex, with irregular objects of diverse shapes and varying scales. These challenges are amplified by heterogeneity and scale disparities across RS data sources, including sensors, platforms, and modalities, making accurate object segmentation particularly demanding. While the Segment Anything Model (SAM) has shown significant potential in segmenting complex scenes, its performance in handling form-varying objects remains limited due to manual-interactive prompting. To this end, we propose UrbanSAM, a customized version of SAM specifically designed to analyze complex urban environments while tackling scaling effects from remotely sensed observations. Inspired by multi-resolution analysis (MRA) theory, UrbanSAM incorporates a novel learnable prompter equipped with a Uscaling-Adapter that adheres to the invariance criterion, enabling the model to capture multiscale contextual information of objects and adapt to arbitrary scale variations with theoretical guarantees. Furthermore, features from the Uscaling-Adapter and the trunk encoder are aligned through a masked cross-attention operation, allowing the trunk encoder to inherit the adapter's multiscale aggregation capability. This synergy enhances the segmentation performance, resulting in more powerful and accurate outputs, supported by the learned adapter. Extensive experimental results demonstrate the flexibility and superior segmentation performance of the proposed UrbanSAM on a global-scale dataset, encompassing scale-varying urban objects such as buildings, roads, and water.

遥感分割多尺度SAM改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。