arXiv:2602.08206cs.CV2026-02

提出地理推理框架,让遥感图像自动理解复杂场景语义。

Geospatial-Reasoning-Driven Vocabulary-Agnostic Remote Sensing Semantic Segmentation

  • 用地理推理链构建类别解释标准,区分易混淆地物
  • 在LoveDA和GID5上提升整体分割性能,复杂场景更准确
  • 适合需要跨类别识别的遥感分析人员

开放词汇语义分割已成为遥感领域的重要方向,可识别预定义类别之外的地物。但现有方法多依赖被动视觉-文本匹配,在地理复杂场景中常因语义模糊而表现不佳,尤其当不同类别具有相似光谱或结构特征时。为此,本文提出地理推理链式思维(GR-CoT)框架,包含离线知识蒸馏流与在线实例推理流。前者为易混淆类别构建分类解释标准,后者实现宏观场景锚定、视觉特征解耦与知识驱动决策融合,生成适配图像的自适应词汇用于下游分割。在LoveDA与GID5基准上的实验表明,该框架显著提升整体分割性能,并在复杂场景中产生更语义一致的预测结果。

原文摘要 · Abstract (English)

Open-vocabulary semantic segmentation has become an important direction in remote sensing, as it enables recognition beyond predefined land-cover categories. However, existing methods mainly depend on passive visual-text matching and often struggle with semantic ambiguity in geographically complex scenes, especially when different classes exhibit similar spectral or structural patterns. To address this issue, we propose a Geospatial Reasoning Chain-of-Thought (GR-CoT) framework for remote sensing open-vocabulary semantic segmentation. GR-CoT consists of an offline knowledge distillation stream and an online instance reasoning stream. The former constructs category interpretation standards for confusing classes, while the latter performs macro-scenario anchoring, visual feature decoupling, and knowledge-driven decision synthesis to generate an image-adaptive vocabulary for downstream segmentation. Experiments on the LoveDA and GID5 benchmarks indicate that the proposed framework improves overall segmentation performance and yields more semantically coherent predictions in complex scenes.

遥感分割开放词汇地理推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。