arXiv:2508.18067cs.CV2025-08被引 10

无需标注即可分割遥感图像中的新类别,突破传统方法依赖人工标注的瓶颈。

Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images

  • 提出SimFeatUp通用上采样器,从低分辨率特征恢复高精度空间细节。
  • 引入全局偏置消除机制,提升局部语义准确性,无需额外训练。
  • 支持光学与雷达遥感多模态统一分割,适用于无基础模型的场景。

遥感图像语义分割对全面地球观测至关重要,但识别新物体类别与高昂的人工标注成本构成重大挑战。现有针对自然图像的开放词汇分割方法难以应对遥感数据的大尺度变化和细粒度特征,且常依赖大量标注。为此,本文提出首个无需标注的遥感开放词汇分割框架SegEarth-OV。我们设计SimFeatUp通用上采样器,可从粗略特征中稳健恢复高分辨率空间细节,无需任务特定后训练即可纠正目标形状失真;提出简单有效的全局偏置消除操作,从块特征中移除固有全局上下文,显著提升局部语义保真度。这些组件使SegEarth-OV能有效利用预训练视觉语言模型(VLM)的丰富语义,在光学遥感中实现开放词汇分割。为进一步拓展至如合成孔径雷达(SAR)等复杂遥感模态,当缺乏大规模VLM且构建成本高昂时,我们提出AlignEarth,一种基于知识蒸馏的策略,可将光学VLM编码器的语义知识高效迁移至SAR编码器,避免从零构建SAR基础模型,实现跨传感器类型的通用开放词汇分割。在光学与SAR数据集上的大量实验表明,SegEarth-OV相较当前最先进方法取得显著提升,为无需标注、开放世界的地球观测奠定坚实基础。

原文摘要 · Abstract (English)

Semantic segmentation of remote sensing (RS) images is pivotal for comprehensive Earth observation, but the demand for interpreting new object categories, coupled with the high expense of manual annotation, poses significant challenges. Although open-vocabulary semantic segmentation (OVSS) offers a promising solution, existing frameworks designed for natural images are insufficient for the unique complexities of RS data. They struggle with vast scale variations and fine-grained details, and their adaptation often relies on extensive, costly annotations. To address this critical gap, this paper introduces SegEarth-OV, the first framework for annotation-free open-vocabulary segmentation of RS images. Specifically, we propose SimFeatUp, a universal upsampler that robustly restores high-resolution spatial details from coarse features, correcting distorted target shapes without any task-specific post-training. We also present a simple yet effective Global Bias Alleviation operation to subtract the inherent global context from patch features, significantly enhancing local semantic fidelity. These components empower SegEarth-OV to effectively harness the rich semantics of pre-trained VLMs, making OVSS possible in optical RS contexts. Furthermore, to extend the framework's universality to other challenging RS modalities like SAR images, where large-scale VLMs are unavailable and expensive to create, we introduce AlignEarth, which is a distillation-based strategy and can efficiently transfer semantic knowledge from an optical VLM encoder to an SAR encoder, bypassing the need to build SAR foundation models from scratch and enabling universal OVSS across diverse sensor types. Extensive experiments on both optical and SAR datasets validate that SegEarth-OV can achieve dramatic improvements over the SOTA methods, establishing a robust foundation for annotation-free and open-world Earth observation.

遥感分割开放词汇无标注多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。