用结构引导提升遥感图像开放词汇分割的跨域泛化能力
GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

- 分离视觉与文本匹配,用预训练模型提供结构指导
- 在多数据集上提升2.5~2.7点平均mIoU,零样本跨区域有效
- 适合需要跨地域、无标注数据的遥感智能解译场景
开放词汇遥感分割能识别训练中未见的类别,但不同地区、分辨率和平台导致的地理域偏移会削弱视觉-文本匹配,限制跨数据集泛化。现有方法虽引入辅助视觉基础模型(VFMs)特征增强匹配,但易引入不一致信号且未充分利用其结构敏感特性。为此,我们提出GeoSeg-OV,将辅助VFMs特征从视觉-文本匹配中解耦,转而作为结构引导用于代价聚合与解码。该方法基于多旋转CLIP特征构建方向鲁棒的代价体,同时冻结的VFMs并行提取多尺度结构敏感特征。提出结构引导聚合(SGA),融合代价令牌、CLIP语义引导与VFMs生成的成对结构偏差,实现连贯的空间传播,并进行文本条件下的类别推理。进一步设计代价感知解码(CAD),根据当前解码上下文自适应融合多尺度语义与结构引导。在覆盖六大洲七数据集的全球高分辨率土地覆盖(HRLC)基准上,两种训练设置下均优于当前最优方法2.5和2.7个百分点平均mIoU。大规模零样本实验验证了其在无目标域标注或重训练情况下的跨地理域与类别系统泛化能力。
原文摘要 · Abstract (English)
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization. Recent attempts have begun to incorporate auxiliary vision foundation models (VFMs), typically coupling their features with text embeddings as additional matching evidence. However, this strategy may introduce inconsistent matching signals while leaving the structure-sensitive representations of VFMs insufficiently exploited. We therefore propose GeoSeg-OV, which decouples auxiliary VFM features from visual-text matching and repurposes them as structural guidance for cost aggregation and decoding. GeoSeg-OV constructs an orientation-robust cost volume from multi-rotation CLIP features, while a frozen VFM extracts multi-scale structure-sensitive features in parallel. We propose Structure-Guided Aggregation (SGA), which integrates cost tokens and CLIP semantic guidance with VFM-derived pairwise structural biases for coherent spatial propagation, followed by text-conditioned class-wise reasoning. We further introduce Cost-Aware Decoding (CAD) to adaptively refine and fuse multi-scale semantic and structural guidance based on the current decoder context. On the global High-Resolution Land Cover (HRLC) benchmark spanning seven datasets across six continents, GeoSeg-OV outperforms the state-of-the-art by +2.5 and +2.7 average mIoU under two training settings. A large-scale zero-shot case study further demonstrates its generalization across geographic domains and category systems without target-domain annotations or retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。