arXiv:2509.12040cs.CVcs.AI2025-09AAAI被引 52

针对遥感图像开放词汇分割,提出高效新框架并建立统一评测基准。

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

  • 设计多方向代价图聚合与轻量融合机制,提升旋转不变性与效率。
  • 在新基准上比基线模型高3.8 mIoU和5.9 mACC,推理速度提升2倍。
  • 适合遥感领域研究者,尤其关注开放词汇分割与跨域适配场景。

开放词汇遥感图像分割(OVRSIS)是新兴任务,因缺乏统一评估基准及自然图像与遥感图像之间的领域差异而研究不足。为此,我们首先基于常用遥感分割数据集构建标准化评测基准(OVRSISBench),实现方法间一致评估。基于此基准,系统评估多个代表性模型,揭示其直接应用于遥感场景的局限性。在此基础上,提出专为遥感设计的新框架RSKT-Seg,包含三个核心模块:(1)多方向代价图聚合(RS-CMA),通过多方向视觉-语言余弦相似性捕捉旋转不变特征;(2)高效代价图融合(RS-Fusion)Transformer,结合轻量降维策略联合建模空间与语义依赖;(3)遥感知识迁移(RS-Transfer)模块,通过增强上采样注入预训练知识,促进域适应。大量实验表明,RSKT-Seg在该基准上持续优于强基线模型,提升3.8 mIoU和5.9 mACC,同时推理速度加快2倍。代码已开源。

原文摘要 · Abstract (English)

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these gaps, we first establish a standardized OVRSIS benchmark (\textbf{OVRSISBench}) based on widely-used RS segmentation datasets, enabling consistent evaluation across methods. Using this benchmark, we comprehensively evaluate several representative OVS/OVRSIS models and reveal their limitations when directly applied to remote sensing scenarios. Building on these insights, we propose \textbf{RSKT-Seg}, a novel open-vocabulary segmentation framework tailored for remote sensing. RSKT-Seg integrates three key components: (1) a Multi-Directional Cost Map Aggregation (RS-CMA) module that captures rotation-invariant visual cues by computing vision-language cosine similarities across multiple directions; (2) an Efficient Cost Map Fusion (RS-Fusion) transformer, which jointly models spatial and semantic dependencies with a lightweight dimensionality reduction strategy; and (3) a Remote Sensing Knowledge Transfer (RS-Transfer) module that injects pre-trained knowledge and facilitates domain adaptation via enhanced upsampling. Extensive experiments on the benchmark show that RSKT-Seg consistently outperforms strong OVS baselines by +3.8 mIoU and +5.9 mACC, while achieving 2x faster inference through efficient aggregation. Our code is \href{https://github.com/LiBingyu01/RSKT-Seg}{\textcolor{blue}{here}}.

遥感分割开放词汇高效模型知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。