arXiv:2605.24532cs.CV2026-05

提出ICIPNet提升遥感图像指代分割的跨模态对齐精度

Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation

论文配图:Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation
图 1 · 摘自论文原文
  • 用图像条件提示模块生成自适应视觉语义表示
  • 双路信息融合提升令牌与通道维度特征对齐
  • 适合遥感目标精准定位与跨模态理解任务

指代遥感图像分割(RRSIS)是与具身感知范式相关的场景化、任务驱动型跨模态任务,要求模型将视觉空间特征与语言意图对齐以实现精准目标感知。近期研究聚焦于细化文本特征粒度并优化图像-文本特征融合,以更好引导目标特征表示。然而,描述粒度不足及对语义变化敏感等问题仍导致跨模态特征融合瓶颈。为此,我们提出图像条件实例提示网络(ICIPNet)与双边信息融合(BIF),以缓解跨模态融合瓶颈。ICIPNet引入图像条件实例提示(ICIP)模块,无需外部知识即可生成自适应视觉与语义表示;BIF模块在令牌与通道维度增强特征融合。实验表明,所提ICIPNet优于现有RRSIS模型。

原文摘要 · Abstract (English)

Referring Remote Sensing Image Segmentation (RRSIS) is a situated, task-driven cross-modal task related to the embodied perception paradigm, requiring models to align visual-spatial features with linguistic intentions for precise target perception. Recent research has focused on refining the granularity of textual features and optimizing image-text feature fusion to better guide target feature representations. However, insufficient descriptive granularity and sensitivity to semantic shifts can cause bottlenecks in cross-modal feature fusion. To address these issues, we propose the Image-Conditioned Instance Prompt Network (ICIPNet) with Bilateral Information Fusion, which is designed to alleviate bottlenecks in cross-modal feature fusion. ICIPNet introduces an Image-Conditioned Instance Prompt (ICIP) module to generate self-adaptive visual and semantic representations without external knowledge. The Bilateral Information Fusion (BIF) module enhances feature fusion along the token and channel dimensions. Experiments demonstrate that the proposed ICIPNet outperforms existing RRSIS models.

遥感分割跨模态对齐提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。