提升遥感图像中基于语言的目标定位能力
RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
- 通过思维链微调增强模型对位置信息的感知
- 设计距离感知奖励函数,提升定位精度
- 适合需要精准空间推理的遥感分析场景
遥感视觉定位(RSVG)旨在根据自然语言描述,在大规模航空影像中定位目标物体。由于遥感场景空间尺度大、语义模糊,语言描述常依赖位置线索,给多模态大模型的空间推理带来挑战。为此,我们提出一种基于推理引导的位置感知后训练框架RSGround-R1,以逐步增强空间理解能力。首先,利用合成生成的RSVG推理数据进行思维链监督微调(CoT-SFT),建立显式的定位意识;随后,引入新型位置奖励机制,通过连续且距离感知的反馈进行强化微调(RFT),实现精准定位;此外,为缓解推理过程中的定位不一致问题,设计了空间一致性引导优化方案,动态调整策略更新,确保模型稳定收敛。在多个RSVG基准测试上,该模型展现出优异性能与泛化能力。
原文摘要 · Abstract (English)
Remote Sensing Visual Grounding (RSVG) aims to localize target objects in large-scale aerial imagery based on natural language descriptions. Owing to the vast spatial scale and high semantic ambiguity of remote sensing scenes, these descriptions often rely heavily on positional cues, posing unique challenges for Multimodal Large Language Models (MLLMs) in spatial reasoning. To leverage this unique feature, we propose a reasoning-guided, position-aware post-training framework, dubbed \textbf{RSGround-R1}, to progressively enhance spatial understanding. Specifically, we first introduce Chain-of-Thought Supervised Fine-Tuning (CoT-SFT) using synthetically generated RSVG reasoning data to establish explicit position awareness. Reinforcement Fine-Tuning (RFT) is then applied, augmented by our newly designed positional reward that provides continuous and distance-aware guidance toward accurate localization. Moreover, to mitigate incoherent localization behaviors across rollouts, we introduce a spatial consistency guided optimization scheme that dynamically adjusts policy updates based on their spatial coherence, ensuring stable and robust convergence. Extensive experiments on RSVG benchmarks demonstrate superior performance and generalization of our model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。