arXiv:2502.08486cs.CV2025-02被引 13

通过双向对齐提升遥感图像语义分割精度

Referring Remote Sensing Image Segmentation via Bidirectional Alignment Guided Joint Prediction

  • 引入双向空间相关性增强视觉与语言特征对齐
  • 在两个基准数据集上实现oIoU提升3.76和1.44个百分点
  • 适合遥感监测、城市规划等需要精准目标识别的场景

遥感图像指代分割(RRSIS)在生态监测、城市规划和灾害管理中至关重要,需根据文本描述精确分割遥感影像中的目标。该任务因视觉-语言差距大、遥感图像分辨率高、覆盖范围广、类别多样且目标细小,以及目标簇集、边缘模糊等问题而极具挑战。为此,本文提出 extit{ours}框架,旨在缩小视觉-语言差距、增强多尺度特征交互、提升细粒度目标区分能力。具体包括:(1) 双向空间相关性(BSC)以改善视觉-语言特征对齐;(2) 目标-背景双流解码器(T-BTD)实现目标与非目标的精确区分;(3) 双模态目标学习策略(D-MOLS)用于鲁棒的多模态特征重建。在基准数据集RefSegRS和RRSIS-D上的大量实验表明, extit{ours}达到当前最优性能:在两个数据集上总体交并比(oIoU)分别提升3.76(80.57)和1.44(79.23)个百分点;平均交并比(mIoU)分别提升5.37(67.95)和1.84(66.04)个百分点,有效解决RRSIS核心挑战,显著提升精度与鲁棒性。

原文摘要 · Abstract (English)

Referring Remote Sensing Image Segmentation (RRSIS) is critical for ecological monitoring, urban planning, and disaster management, requiring precise segmentation of objects in remote sensing imagery guided by textual descriptions. This task is uniquely challenging due to the considerable vision-language gap, the high spatial resolution and broad coverage of remote sensing imagery with diverse categories and small targets, and the presence of clustered, unclear targets with blurred edges. To tackle these issues, we propose \ours, a novel framework designed to bridge the vision-language gap, enhance multi-scale feature interaction, and improve fine-grained object differentiation. Specifically, \ours introduces: (1) the Bidirectional Spatial Correlation (BSC) for improved vision-language feature alignment, (2) the Target-Background TwinStream Decoder (T-BTD) for precise distinction between targets and non-targets, and (3) the Dual-Modal Object Learning Strategy (D-MOLS) for robust multimodal feature reconstruction. Extensive experiments on the benchmark datasets RefSegRS and RRSIS-D demonstrate that \ours achieves state-of-the-art performance. Specifically, \ours improves the overall IoU (oIoU) by 3.76 percentage points (80.57) and 1.44 percentage points (79.23) on the two datasets, respectively. Additionally, it outperforms previous methods in the mean IoU (mIoU) by 5.37 percentage points (67.95) and 1.84 percentage points (66.04), effectively addressing the core challenges of RRSIS with enhanced precision and robustness.

遥感分割图文对齐多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。