arXiv:2604.01893cs.CV2026-04被引 5

通过解耦语言线索,实现遥感图像的渐进式精确定位

ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery

  • 将语言描述拆分为全局上下文、空间关系和对象属性,分阶段引导定位
  • 在两个基准上达到新最好性能,定位准确率显著提升
  • 适合需要细粒度语义理解的遥感图像分析任务

遥感视觉定位(RSVG)旨在根据自然语言描述定位遥感图像中的目标。现有方法多依赖句级视觉-语言对齐,难以利用空间关系和对象属性等细粒度语言线索,而这些线索对区分特征相似的目标至关重要。本文提出ProVG框架,通过解耦语言表达为全局上下文、空间关系和对象属性,分阶段提供更明确的指导。采用简单的渐进式跨模态调制器,基于‘概览-定位-验证’流程动态调节视觉注意力,实现从粗到精的对齐。同时引入跨尺度融合模块缓解遥感图像的大尺度差异,并设计语言引导校准解码器,在预测中优化跨模态对齐。统一的多任务头支持指代表达理解和分割任务。在RRSIS-D和RISBench两个基准上的大量实验表明,ProVG持续优于现有方法,达到新的最先进水平。

原文摘要 · Abstract (English)

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, which struggles to exploit fine-grained linguistic cues, such as \textit{spatial relations} and \textit{object attributes}, that are crucial for distinguishing objects with similar characteristics. Importantly, these cues play distinct roles across different grounding stages and should be leveraged accordingly to provide more explicit guidance. In this work, we propose \textbf{ProVG}, a novel RSVG framework that improves localization accuracy by decoupling language expressions into global context, spatial relations, and object attributes. To integrate these linguistic cues, ProVG employs a simple yet effective progressive cross-modal modulator, which dynamically modulates visual attention through a \textit{survey-locate-verify} scheme, enabling coarse-to-fine vision-language alignment. In addition, ProVG incorporates a cross-scale fusion module to mitigate the large-scale variations in remote sensing imagery, along with a language-guided calibration decoder to refine cross-modal alignment during prediction. A unified multi-task head further enables ProVG to support both referring expression comprehension and segmentation tasks. Extensive experiments on two benchmarks, \textit{i.e.}, RRSIS-D and RISBench, demonstrate that ProVG consistently outperforms existing methods, achieving new state-of-the-art performance.

遥感图像视觉定位语言解耦多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。