arXiv:2505.00738eess.IVcs.LG2025-05

提出细粒度遥感场景文本定位任务,精准识别中尺度语义区域。

XeMap: Contextual Referring in Large-Scale Remote Sensing Environments

  • 设计融合自注意力与交叉注意力的网络架构,增强图文交互。
  • 引入分层多尺度语义对齐模块,实现跨尺度精准匹配。
  • 构建专用数据集XeMap-set,支持零样本测试,性能领先。

遥感影像技术的发展带来了高分辨率与大范围覆盖,但现有图像级描述/检索和目标级检测/分割方法难以捕捉大规模场景中的中尺度语义实体。为此,我们提出上下文性文本定位任务(XeMap),聚焦于在大尺度遥感场景中对文本所指区域进行细粒度定位。不同于传统方法,XeMap能精确映射常被忽略的中尺度语义实体。为此,我们提出XeMap-Network,一种用于像素级跨模态上下文定位的新型架构,包含融合层,采用自注意力与交叉注意力机制提升文本与图像嵌入的交互能力;同时提出分层多尺度语义对齐(HMSA)模块,将多尺度视觉特征与文本语义向量对齐,实现大尺度遥感图像中的精确多模态匹配。为支持该任务,我们构建了专用于此任务的新标注数据集XeMap-set,解决了遥感领域缺乏此类数据集的问题。XeMap-Network在零样本设置下对比前沿方法表现更优,验证了其在精准定位文本所指区域上的有效性,为理解大尺度遥感环境提供了新视角。

原文摘要 · Abstract (English)

Advancements in remote sensing (RS) imagery have provided high-resolution detail and vast coverage, yet existing methods, such as image-level captioning/retrieval and object-level detection/segmentation, often fail to capture mid-scale semantic entities essential for interpreting large-scale scenes. To address this, we propose the conteXtual referring Map (XeMap) task, which focuses on contextual, fine-grained localization of text-referred regions in large-scale RS scenes. Unlike traditional approaches, XeMap enables precise mapping of mid-scale semantic entities that are often overlooked in image-level or object-level methods. To achieve this, we introduce XeMap-Network, a novel architecture designed to handle the complexities of pixel-level cross-modal contextual referring mapping in RS. The network includes a fusion layer that applies self- and cross-attention mechanisms to enhance the interaction between text and image embeddings. Furthermore, we propose a Hierarchical Multi-Scale Semantic Alignment (HMSA) module that aligns multiscale visual features with the text semantic vector, enabling precise multimodal matching across large-scale RS imagery. To support XeMap task, we provide a novel, annotated dataset, XeMap-set, specifically tailored for this task, overcoming the lack of XeMap datasets in RS imagery. XeMap-Network is evaluated in a zero-shot setting against state-of-the-art methods, demonstrating superior performance. This highlights its effectiveness in accurately mapping referring regions and providing valuable insights for interpreting large-scale RS environments.

遥感文本定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。