arXiv:2606.28920cs.CV2026-06

无需训练,用一张示例图精准定位遥感图像中的目标

ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images

论文配图:ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images
图 1 · 摘自论文原文
  • 用一张示例图提供结构指引,修正语言模型的粗略定位
  • 通过迭代聚类生成高质量正负样本提示,提升边界精度
  • 适合遥感图像中无标注数据的精准目标定位任务

遥感视觉定位(RSVG)旨在利用自由格式自然语言描述,在高分辨率遥感图像中定位特定目标。尽管多模态大语言模型(MLLMs)在开放词汇的RSVG中展现出巨大潜力,但其无训练适配受制于抽象语义与细粒度视觉线索之间的模态差距。在复杂遥感场景中,该差距导致严重定位漂移。为此,我们提出示例驱动的校准精炼框架ExACT,采用单次视觉提示机制,为像素级定位提供明确的结构引导。具体地,提出视觉示例校准器(VEC),从给定示例中提取细粒度视觉对应关系,以修正冻结MLLM产生的粗略跨模态先验,有效抑制背景噪声并精确勾勒目标边界。随后,结构感知精炼器(SAR)采用迭代合并-选择聚类策略,将校准后的先验整合为高质量正负几何提示。这些提示引导分割任意模型(SAM)实现精确像素级预测。大量实验验证了ExACT在无训练与弱监督方法中的优越性。

原文摘要 · Abstract (English)

Remote sensing visual grounding (RSVG) aims to locate specific objects in high-resolution RS imagery using free-form natural language descriptions. While recent advances in multimodal large language models (MLLMs) show great potential for such open-vocabulary RSVG, their training-free adaptation is hindered by the modality gap between abstract linguistic semantics and fine-grained visual cues. In cluttered RS scenes, this gap inevitably causes severe localization drift. To bridge this gap, we propose Exemplar-driven Calibrated Refinement (ExACT), a novel training-free framework driven by a one-shot visual prompting mechanism to explicitly provide discriminative structural guidance for precise pixel-level localization. Specifically, we propose a Vision Exemplar-based Calibrator (VEC) that extracts fine-grained visual correspondences from the given exemplar to rectify the rough cross-modal priors from frozen MLLMs, effectively suppressing background artifacts and accurately outlining target boundaries. Subsequently, a Structure-Aware Refiner (SAR) employs an iterative merge-and-select clustering strategy to consolidate the calibrated priors into high-quality positive and negative geometric prompts. These prompts then guide the Segment Anything Model (SAM) to achieve precise pixel-level predictions. Extensive experiments confirm the superiority of ExACT over existing training-free and weakly-supervised methods.

遥感图像视觉定位无训练提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。