arXiv:2507.19280cs.CV2025-07AAAI被引 28

用强化学习让遥感模型自主推理,统一处理多粒度空间任务。

RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow

  • 基于多模态大模型与强化学习,实现自主推理路径探索。
  • 在多粒度任务上达到当前最佳性能,无需额外微调。
  • 适合需要通用遥感理解的科研与应用开发者。

遥感影像包含大量结构化程度低的空间数据,需复杂推理才能理解用户意图与上下文关系,超越简单识别任务。本文旨在构建一个统一的地球观测推理工作流,以应对复杂查询。该工作流应能自主探索并构建推理路径,而非局限于预定义的固定流程。理想架构需统一且通用,通过单一模型完成多种推理任务,无需额外微调。现有方法依赖监督微调和特定任务头,限制了自主推理与统一泛化能力。为此,我们提出RemoteReasoner,一种统一的地理空间推理工作流。其设计融合多模态大语言模型(MLLM)以解析用户指令并定位目标,结合任务转换策略支持对象、区域与像素级等多粒度任务。与现有方法不同,本框架采用强化学习训练,赋予MLLM充分的推理自主性。推理阶段,转换策略可生成多样化输出格式,无需特定解码器或再微调。实验表明,RemoteReasoner在多粒度推理任务上达到当前最优(SOTA)表现,同时保持MLLM固有的泛化能力,在未见任务与分布外类别上仍表现稳健。

原文摘要 · Abstract (English)

Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to construct an Earth observation workflow to handle complex queries by reasoning about spatial context and user intent. As a reasoning workflow, it should autonomously explore and construct its own inference paths, rather than being confined to predefined ground-truth sequences. Ideally, its architecture ought to be unified yet generalized, possessing capabilities to perform diverse reasoning tasks through one model without requiring additional fine-tuning. Existing remote sensing approaches rely on supervised fine-tuning paradigms and task-specific heads, limiting both autonomous reasoning and unified generalization. To this end, we propose RemoteReasoner, a unified workflow for geospatial reasoning. The design of RemoteReasoner integrates a multi-modal large language model (MLLM) for interpreting user instructions and localizing targets, together with task transformation strategies that enable multi-granularity tasks, including object-, region-, and pixel-level. In contrast to existing methods, our framework is trained with reinforcement learning (RL) to endow the MLLM sufficient reasoning autonomy. At the inference stage, our transformation strategies enable diverse task output formats without requiring task-specific decoders or further fine-tuning. Experiments demonstrated that RemoteReasoner achieves state-of-the-art (SOTA) performance across multi-granularity reasoning tasks. Furthermore, it retains the MLLM's inherent generalization capability, demonstrating robust performance on unseen tasks and out-of-distribution categories.

遥感多模态强化学习推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。