arXiv:2512.08881cs.CV2025-12

提升卫星图像视觉定位精度,通过空间感知机制让模型更准找目标

SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing

  • 用控制令牌接入专用定位模块,实现语言与空间信息联合推理
  • 在遥感基准上实现33.2%相对提升,显著优于现有方法
  • 适合需要精准定位的卫星图像分析场景,如灾害监测、城市规划

视觉-语言模型(VLMs)正成为遥感领域的强大通用工具,可整合多任务信息并支持基于指令的交互。本文提出一种新型结构化定位机制,通过微调预训练VLM完成多样化指令任务,并引入专用控制令牌连接独立定位模块。该方法使模型能联合推理语言与空间信息,在复杂卫星场景中显著提升目标定位精度。我们在多个遥感基准上评估,持续超越当前最优水平,尤其在视觉定位任务上相较先前方法实现33.2%的相对性能提升。结果表明,将结构化空间推理融入VLM具有显著优势,为更可靠的卫星数据实际应用铺平道路。代码将在论文接受后公开。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are emerging as powerful generalist tools for remote sensing, capable of integrating information across diverse tasks and enabling flexible, instruction-based interactions via a chat interface. In this work, we enhance VLM-based visual grounding in satellite imagery by proposing a novel structured localization mechanism. Our approach involves finetuning a pretrained VLM on a diverse set of instruction-following tasks, while interfacing a dedicated grounding module through specialized control tokens for localization. This method facilitates joint reasoning over both language and spatial information, significantly enhancing the model's ability to precisely localize objects in complex satellite scenes. We evaluate our framework on several remote sensing benchmarks, consistently improving the state-of-the-art, including a 33.2% relative improvement over previous methods on visual grounding. Our results highlight the benefits of integrating structured spatial reasoning into VLMs, paving the way for more reliable real-world satellite data analysis. Code will be released upon acceptance.

遥感视觉定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。