arXiv:2509.21976cs.CVcs.AI2025-09中稿 · ISPRS被引 20

用强化学习让模型先推理再定位,提升少样本遥感指代表达理解能力

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

  • 先生成可解释的推理链,再据此定位目标对象
  • 在三个少样本遥感基准上显著优于监督微调方法
  • 适合遥感、地理信息等数据稀缺场景下的模型开发

遥感中的指代表达理解面临独特挑战,需对复杂物体-上下文关系进行推理。尽管基于大规模标注数据的多模态大模型监督微调表现优异,但在数据稀缺场景下泛化能力差。为此,我们提出Geo-R1,一种以推理为核心的强化微调(RFT)范式,用于少样本遥感指代表达理解。Geo-R1强制模型先生成显式的、可解释的推理链,分解指代表达,再利用这些推理链定位目标物体。这一“先推理,后行动”流程使模型更有效地利用有限标注,提升泛化能力并增强可解释性。我们在三个精心设计的少样本遥感指代表达基准上验证了Geo-R1,结果表明其性能持续且显著优于SFT基线,并展现出强跨数据集泛化能力。代码与数据将公开于:https://github.com/Geo-R1/geo-r1。

原文摘要 · Abstract (English)

Referring expression understanding in remote sensing poses unique challenges, as it requires reasoning over complex object-context relationships. While supervised fine-tuning (SFT) on multimodal large language models achieves strong performance with massive labeled datasets, they struggle in data-scarce scenarios, leading to poor generalization. To address this limitation, we propose Geo-R1, a reasoning-centric reinforcement fine-tuning (RFT) paradigm for few-shot geospatial referring. Geo-R1 enforces the model to first generate explicit, interpretable reasoning chains that decompose referring expressions, and then leverage these rationales to localize target objects. This "reason first, then act" process enables the model to make more effective use of limited annotations, enhances generalization, and provides interpretability. We validate Geo-R1 on three carefully designed few-shot geospatial referring benchmarks, where our model consistently and substantially outperforms SFT baselines. It also demonstrates strong cross-dataset generalization, highlighting its robustness. Code and data will be released at: https://github.com/Geo-R1/geo-r1.

遥感理解少样本学习强化微调可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。