arXiv:2606.16124cs.CV2026-06

无需训练即可精准定位遥感图像中的任意语义目标

Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

论文配图:Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos
图 1 · 摘自论文原文
  • 用视觉语言模型和扩散模型协同生成注意力图,逐步优化定位结果
  • 在6个基准上超越零样本基线,性能媲美有监督方法
  • 适合需要快速适配新场景的遥感智能分析任务

遥感视觉定位(RSVG)旨在根据自然语言描述定位遥感图像或视频中的目标。现有方法依赖特定任务的人工标注,成本高且难以覆盖真实地理场景多样性,导致对新物体、细粒度属性、复杂空间关系和功能语义等开放词汇查询泛化能力差。本文提出 RSVG-ZeroOV,一种无需训练的框架,利用冻结的通用基础模型实现零样本开放词汇遥感视觉定位。该框架采用“概览-聚焦-演化”范式,充分利用视觉语言模型(VLM)与扩散模型(DM)在注意力模式上的互补性,逐步生成精确定位结果:(i) 概览阶段使用VLM提取跨注意力图,捕捉语言表达与视觉区域间的语义关联;(ii) 聚焦阶段借助DM的细粒度建模先验,补足VLM忽略的对象结构与形状信息;(iii) 演化阶段引入简单有效的注意力演化模块,抑制无关激活,生成纯净的目标掩码。针对视频输入,进一步提出 Video RSVG-ZeroOV,通过查询相关关键帧选择器与时间传播器,将图像级定位扩展为时空定位,实现无需视频标注或微调的高效、时序一致视频定位。在六个图像与视频定位基准上的大量实验表明,RSVG-ZeroOV持续优于现有零样本基线,并达到与弱监督及全监督方法相当或更优的性能。

原文摘要 · Abstract (English)

Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usually rely on task-specific manual annotations, which are costly to collect and inevitably limited in covering the diversity of real-world geospatial scenarios. As a result, they often struggle to generalize to open-vocabulary queries involving novel objects, fine-grained attributes, complex spatial relationships, and functional semantics. In this paper, we propose RSVG-ZeroOV, a training-free framework that leverages frozen generic foundation models for zero-shot open-vocabulary RSVG. RSVG-ZeroOV follows an Overview-Focus-Evolve paradigm, which exploits the distinct yet complementary attention patterns of vision-language models (VLMs) and diffusion models (DMs) to progressively generate precise grounding results. Specifically, (i) Overview utilizes a VLM to extract cross-attention maps that capture semantic correlations between the referring expression and visual regions; (ii) Focus leverages the fine-grained modeling priors of a DM to compensate for object structure and shape information often overlooked by VLM attention; and (iii) Evolve introduces a simple yet effective attention evolution module to suppress irrelevant activations, yielding purified object masks. To handle video inputs, we further present Video RSVG-ZeroOV, which extends image-level grounding to spatio-temporal grounding through a query-relevant key-frame selector and a temporal propagator, enabling efficient and temporally coherent video grounding without video annotations or fine-tuning. Extensive experiments on six image and video grounding benchmarks show that RSVG-ZeroOV consistently outperforms existing zero-shot baselines and achieves competitive or superior performance compared with weakly- and fully-supervised methods.

遥感定位零样本视觉语言模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。