arXiv:2512.19302cs.CV2025-12被引 3

让大模型精准指挥分割模型,实现遥感图像语义与几何的精准结合。

Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing

  • 解耦语言推理与分割,用强化学习让大模型生成空间引导提示。
  • 在EarthReason数据集上达到75.60% cIoU和73.36% gIoU,超越最强基线。
  • 适合需要精准语义分割的遥感分析任务,尤其关注跨任务泛化能力。

大型视觉-语言模型(LVLM)在光学遥感(RS)分析中潜力巨大,但现有推理分割框架通过端到端微调将语言推理与像素预测耦合,导致几何定位能力弱、任务泛化性差。为此,我们提出Think2Seg-RS,一种解耦框架:训练一个LVLM提示器,通过结构化几何提示控制冻结的Segment Anything Model(SAM)。采用仅基于最终掩码IoU的组相对策略优化(GRPO)强化学习目标,使LVLM学会将抽象语义推理转化为空间可落地的动作,在EarthReason数据集上达到领先性能。显著地,Think2Seg-RS在EarthReason上测试cIoU达75.60%,gIoU达73.36%,相较最强基线分别提升6.47%和2.40%。零样本评估在三个指代分割基准上揭示任务归纳偏置的根本差异:语义级定位(聚合所有符合概念意图区域)与实例级任务(需分离离散对象)存在本质区别。我们进一步发现,在语义级监督下,紧凑分割器优于大模型,能缓解纹理过分割;而无约束负向提示在异质航拍背景中不稳定。这些结果表明,通过直接分割反馈优化LVLM,可构建可扩展的复杂地理空间推理框架,有效弥合抽象语言理解与像素级执行之间的鸿沟。

原文摘要 · Abstract (English)

Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation frameworks couple linguistic reasoning and pixel prediction through end-to-end supervised fine-tuning, leading to weak geometric grounding and limited generalization across tasks. To address this, we developed Think2Seg-RS, a decoupled framework that trains an LVLM prompter to control a frozen Segment Anything Model (SAM) via structured geometric prompts. Through a mask-only Group Relative Policy Optimization (GRPO) reinforcement learning objective driven strictly by final mask IoU, the LVLM learns to translate abstract semantic reasoning into spatially grounded actions, achieving state-of-the-art performance on the EarthReason dataset. Notably, Think2Seg-RS outperforms leading approaches such as RemoteReasoner and SegEarth-R1 on the EarthReason dataset by reaching a test cIoU of 75.60% and gIoU of 73.36%, yielding absolute improvements of 6.47% and 2.40% over the strongest baseline, respectively. Zero-shot evaluations across three referring segmentation benchmarks reveal a fundamental distinction in task inductive bias, exposing a distinct divide between semantic-level grounding -- which aggregates all regions matching a conceptual intent -- and instance-level tasks that demand discrete object separation. We further found that compact segmenters outperform larger ones under semantic-level supervision by mitigating textural over-segmentation, and that unconstrained negative prompting is unstable in heterogeneous aerial backgrounds. Together, these findings demonstrate that optimizing LVLMs through direct segmentation feedback offers a scalable framework for complex geospatial reasoning, effectively bridging the gap between abstract language understanding and precise pixel-level execution.

遥感分割大模型语义几何强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。