arXiv:2410.08613cs.CVcs.AI2024-10被引 36

提出跨模态双向交互模型,精准分割遥感图像中语言描述的目标

Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation

  • 设计双向交互模块,融合语言与多尺度视觉特征
  • 在52,472样本数据集上超越现有最佳方法
  • 适合遥感图像语义理解与地理信息分析研究者

给定自然语言表达和遥感图像,遥感图像指代分割(RRSIS)旨在生成由语言描述所指目标的像素级掩码。与自然场景相比,遥感场景中的语言表达常包含复杂的地理空间关系,目标物体尺度差异大且缺乏视觉显著性,导致分割难度高。为此,本文提出一种新型框架——跨模态双向交互模型(CroBIM)。设计上下文感知提示调制(CAPM)模块,将空间位置关系和任务知识融入语言特征,增强目标捕捉能力;引入语言引导特征聚合(LGFA)模块,将语言信息注入多尺度视觉特征,并采用注意力补偿机制提升特征聚合效果;设计相互作用解码器(MID),通过级联双向交叉注意力增强跨模态特征对齐,实现精确掩码预测。为进一步推动研究,构建了包含52,472组图像-语言-标签三元组的新基准数据集RISBench。在RISBench及另外两个主流数据集上的大量实验表明,所提CroBIM优于现有最先进方法。代码与数据集将公开于https://github.com/HIT-SIRS/CroBIM。

原文摘要 · Abstract (English)

Given a natural language expression and a remote sensing image, the goal of referring remote sensing image segmentation (RRSIS) is to generate a pixel-level mask of the target object identified by the referring expression. In contrast to natural scenarios, expressions in RRSIS often involve complex geospatial relationships, with target objects of interest that vary significantly in scale and lack visual saliency, thereby increasing the difficulty of achieving precise segmentation. To address the aforementioned challenges, a novel RRSIS framework is proposed, termed the cross-modal bidirectional interaction model (CroBIM). Specifically, a context-aware prompt modulation (CAPM) module is designed to integrate spatial positional relationships and task-specific knowledge into the linguistic features, thereby enhancing the ability to capture the target object. Additionally, a language-guided feature aggregation (LGFA) module is introduced to integrate linguistic information into multi-scale visual features, incorporating an attention deficit compensation mechanism to enhance feature aggregation. Finally, a mutual-interaction decoder (MID) is designed to enhance cross-modal feature alignment through cascaded bidirectional cross-attention, thereby enabling precise segmentation mask prediction. To further forster the research of RRSIS, we also construct RISBench, a new large-scale benchmark dataset comprising 52,472 image-language-label triplets. Extensive benchmarking on RISBench and two other prevalent datasets demonstrates the superior performance of the proposed CroBIM over existing state-of-the-art (SOTA) methods. The source code for CroBIM and the RISBench dataset will be publicly available at https://github.com/HIT-SIRS/CroBIM

遥感分割跨模态语言引导图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。