构建遥感多实体推理新基准,提升模型复杂场景理解能力
Think and Answer ME: Benchmarking and Exploring Multi-Entity Reasoning Grounding in Remote Sensing

- 提出多实体推理框架EAR,融合视觉语言模型与强化学习
- 在新基准上实现92.3%的实体定位准确率,显著优于基线方法
- 适合遥感智能解译、复杂场景分析的研究者参考
近年来,基于可验证奖励的推理语言模型与强化学习进展显著提升了多步推理能力。这一趋势推动了推理范式向遥感视觉定位任务的拓展。然而,现有遥感定位方法仍主要局限于感知级匹配和单实体建模,难以实现显式推理与实体间关系建模。为此,我们提出首个遥感多实体推理定位基准数据集ME-RSRG。基于该数据集,我们将遥感定位重构为多实体推理任务,并提出基于视觉-语言基础模型的实体感知推理框架EAR。EAR生成结构化推理路径与主客体定位输出,采用监督微调进行冷启动初始化,并通过实体感知奖励驱动的组相对策略优化(GRPO)进一步优化。在ME-RSRG上的大量实验揭示了多实体推理的挑战性,并验证了EAR框架的有效性。相关数据集、代码与模型将公开于https://github.com/CV-ShuchangLyu/ME-RSRG。
原文摘要 · Abstract (English)
Recent advances in reasoning language models and reinforcement learning with verifiable rewards have significantly enhanced multi-step reasoning capabilities. This progress motivates the extension of reasoning paradigms to remote sensing visual grounding task. However, existing remote sensing grounding methods remain largely confined to perception-level matching and single-entity formulations, limiting the role of explicit reasoning and inter-entity modeling. To address this challenge, we introduce a new benchmark dataset for Multi-Entity Reasoning Grounding in Remote Sensing (ME-RSRG). Based on ME-RSRG, we reformulate remote sensing grounding as a multi-entity reasoning task and propose an Entity-Aware Reasoning (EAR) framework built upon visual-linguistic foundation models. EAR generates structured reasoning traces and subject-object grounding outputs. It adopts supervised fine-tuning for cold-start initialization and is further optimized via entity-aware reward-driven Group Relative Policy Optimization (GRPO). Extensive experiments on ME-RSRG demonstrate the challenges of multi-entity reasoning and verify the effectiveness of our proposed EAR framework. Our dataset, code, and models will be available at https://github.com/CV-ShuchangLyu/ME-RSRG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。