arXiv:2512.02715cs.CV2025-12被引 3

让遥感图像定位更准:用逐步搜索+空间推理找小目标

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

  • 把遥感定位变成一步步搜索,结合视觉与空间推理
  • 在五个数据集上表现超越现有方法,小目标定位更精准
  • 适合需要精确地理定位的遥感分析人员

多模态大模型在视觉定位任务中取得显著进展,实现了文本查询与图像区域间的细粒度跨模态对齐。然而,将此类能力迁移至遥感图像仍面临挑战:目标通常在千米级场景中极小,且查询常涉及相对位置、空间层级或远距离对象的上下文依赖等复杂地理关系。为此,我们提出GeoViS,一种地理奖励驱动的遥感视觉定位框架,将遥感视觉定位重新建模为渐进式的搜索与推理过程。该框架不直接一步预测目标位置,而是通过树状结构的视觉线索序列主动探索全局图像,融合多模态感知、空间推理与奖励引导的探索机制,迭代优化地理空间假设。该设计使模型既能检测微小目标,又能保持对全场景的感知。在五个遥感定位基准上的大量实验表明,GeoViS实现了精确的地理空间理解,在关键视觉定位指标上持续优于现有方法,展现出强大的跨域泛化能力和可解释性。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models(MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring such capabilities to remote sensing imagery remains challenging, as targets are often extremely small within kilometer-scale scenes, and queries typically involve intricate geospatial relations such as relative positions, spatial hierarchies, or contextual dependencies across distant objects. To address these challenges, we propose GeoViS, a Geospatially Rewarded Visual Search framework that reformulates remote sensing visual grounding as a progressive search-and-reasoning process. Rather than directly predicting the target location in a single step, GeoViS actively explores the global image through a tree-structured sequence of visual cues, integrating multimodal perception, spatial reasoning, and reward-guided exploration to refine geospatial hypotheses iteratively. This design enables the model to detect subtle small-scale targets while maintaining holistic scene awareness. Extensive experiments on five remote sensing grounding benchmarks demonstrate that GeoViS achieves precise geospatial understanding and consistently surpasses existing methods across key visual grounding metrics, highlighting its strong cross-domain generalization and interpretability.

遥感定位空间推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。