arXiv:2606.30244cs.CV2026-06

提出轻量级框架S4ECA,高效对齐遥感图像与语言描述。

Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation

论文配图:Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation
图 1 · 摘自论文原文
  • 设计双编码器适配器,用可学习查询提取语义文本代理。
  • 仅更新2.4%参数,在两个遥感数据集上达顶尖性能。
  • 适合需要高效跨模态对齐的遥感智能分析场景。

指代式遥感图像分割(RRSIS)旨在定位并分割遥感图像中由自然语言表达指定的目标物体或区域。现有模型虽借助大规模基础模型取得进展,但多依赖全量微调,计算开销大且易削弱预训练模型的泛化能力,因在较小下游数据集上的过度微调会破坏预训练阶段学到的结构化特征表示。尽管参数高效微调(PET)是潜在替代方案,但现有框架多聚焦单模态优化,难以捕捉多模态推理所需的复杂跨模态依赖关系,同时难以弥合自然场景与航拍影像间的巨大领域差异。为此,我们提出新框架S4ECA,通过参数高效适应实现有效的跨模态交互。具体而言,设计双编码器适配器:文本适配器采用可学习查询从词级嵌入中提炼高语义语言代理,实现早期定位;视觉适配器通过多尺度密集提取器精炼层次化特征表示,并引入语言引导的尺度与空间选择机制,动态强调相关视觉上下文,确保精确跨模态对齐。仅更新2.4%主干参数,模型在RRSIS-D和RefSegRS数据集上达到当前最优性能,证明其在复杂航拍场景中的卓越效率与精度。

原文摘要 · Abstract (English)

Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression in a remote sensing image. While existing RRSIS models have benefited from large-scale foundation models, they predominantly rely on full fine-tuning. These approaches are computationally intensive and may weaken the generalization ability of pre-trained models, as extensive fine-tuning on significantly smaller downstream datasets can distort the well-structured feature representations learned during large-scale pre-training. Although Parameter-Efficient Tuning (PET) offers a potential alternative, existing PET frameworks primarily focus on single-modal optimization, failing to capture the complex cross-modal dependencies required for multimodal reasoning, while simultaneously struggling to bridge the substantial domain gap between natural scenes and aerial imagery. To address these limitations, we propose a novel framework, Semantic-driven Scale and Spatial Selection for Efficient Cross-modal Alignment (S4ECA), which enables effective and efficient cross-modal interaction through parameter-efficient adaptation. Specifically, we design a dual-encoder adapter architecture. The textual adapter employs learnable queries to distill highly semantic language proxies from word-level embeddings, facilitating early grounding. Simultaneously, the visual adapter refines hierarchical feature representations through a multi-scale dense extractor, followed by a language-guided scale and spatial selection mechanism that dynamically emphasizes relevant visual contexts, ensuring precise cross-modal alignment. By updating only 2.4% of the backbone parameters, our proposed model achieves state-of-the-art performance on the RRSIS-D and RefSegRS datasets, demonstrating superior efficiency and precision in complex aerial scenarios.

遥感分割跨模态对齐参数高效轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。