arXiv:2509.25026cs.CV2025-09被引 16

用强化学习提升遥感图像的推理能力,让AI更懂地球观测任务。

GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning

  • 引入任务感知奖励机制,优化遥感视觉语言模型的推理过程。
  • 在多个遥感基准上超越当前最先进模型,提升稳定性和鲁棒性。
  • 适合遥感分析、地理信息、智能地球观测等领域的研究者使用。

近年来,强化学习(RL)在自然图像领域展现出强大的推理能力,但在地球观测(EO)领域的潜力尚未充分探索。EO任务包含目标检测、图像/区域描述、变化检测、定位和时序分析等独特挑战,需要具备任务感知的推理能力。本文提出一种新型后训练框架,通过引入任务感知奖励,有效适配基于推理的强化学习模型以应对多样化的遥感任务。该策略显著提升了遥感图像的推理能力,稳定了优化过程,并增强了模型鲁棒性。在多个遥感基准上的广泛实验表明,该方法在性能上持续优于当前最先进的通用与专用视觉语言模型。代码与模型将公开发布于 https://mustansarfiaz.github.io/GeoVLM-R1/。

原文摘要 · Abstract (English)

Recent advances in reinforcement learning (RL) have delivered strong reasoning capabilities in natural image domains, yet their potential for Earth Observation (EO) remains largely unexplored. EO tasks introduce unique challenges, spanning referred object detection, image or region captioning, change detection, grounding, and temporal analysis, that demand task aware reasoning. We propose a novel post training framework that incorporates task aware rewards to enable effective adaptation of reasoning based RL models to diverse EO tasks. This training strategy enhances reasoning capabilities for remote sensing images, stabilizes optimization, and improves robustness. Extensive experiments across multiple EO benchmarks show consistent performance gains over state of the art generic and specialized vision language models. Code and models will be released publicly at https://mustansarfiaz.github.io/GeoVLM-R1/ .

遥感强化学习视觉语言模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。