无需训练即可实现遥感图像的指令驱动分割,解决领域数据稀缺难题。
GeoSeg: Training-Free Reasoning-Driven Segmentation in Remote Sensing Imagery
- 利用多模态大模型推理与坐标修正结合,实现零样本定位。
- 在GeoSeg-Bench上超越所有基线方法,最高提升达18.7%。
- 适合遥感图像语义分割、少样本场景下的快速部署应用。
多模态大模型(MLLM)正在将分割任务从固定类别预测转向指令引导定位。尽管自然场景中的基于推理的分割进展迅速,但遥感图像因推理导向数据成本高且存在俯视视角等特有挑战,缺乏通用解决方案。本文提出GeoSeg——一种零样本、无需训练的框架,突破推理驱动遥感分割的监督瓶颈。GeoSeg通过两项核心技术实现精准定位:(i) 偏差感知的坐标精修,校正系统性定位偏差;(ii) 双路提示机制,融合语义意图与细粒度空间线索。同时构建了GeoSeg-Bench,包含810对图像-查询组合,具备多层次难度。实验表明,GeoSeg在各项指标上持续优于所有基线,消融实验证实各模块的有效性与必要性。
原文摘要 · Abstract (English)
Recent advances in MLLMs are reframing segmentation from fixed-category prediction to instruction-grounded localization. While reasoning based segmentation has progressed rapidly in natural scenes, remote sensing lacks a generalizable solution due to the prohibitive cost of reasoning-oriented data and domain-specific challenges like overhead viewpoints. We present GeoSeg, a zero-shot, training-free framework that bypasses the supervision bottleneck for reasoning-driven remote sensing segmentation. GeoSeg couples MLLM reasoning with precise localization via: (i) bias-aware coordinate refinement to correct systematic grounding shifts and (ii) a dual-route prompting mechanism to fuse semantic intent with fine-grained spatial cues. We also introduce GeoSeg-Bench, a diagnostic benchmark of 810 image--query pairs with hierarchical difficulty levels. Experiments show that GeoSeg consistently outperforms all baselines, with extensive ablations confirming the effectiveness and necessity of each component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。