arXiv:2603.17528cs.CV2026-03被引 3

融合光学与雷达图像,实现恶劣天气下的开放词汇遥感分割

MM-OVSeg:Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing

  • 通过多模态对齐与双编码器融合,统一光学与雷达特征
  • 在多种云况下显著提升分割鲁棒性与泛化能力
  • 适合遥感图像分析、气象敏感场景的开放词汇任务

开放词汇分割可实现从开放文本类别中进行像素级识别,突破固定类别的限制。尽管在遥感领域潜力巨大,当前进展仍主要集中于晴天光学图像,在多云或雾霾条件下表现不佳。我们提出MM-OVSeg,一种用于恶劣天气下稳健开放词汇分割的多模态光学-SAR融合框架。该方法利用光学影像丰富的光谱语义和合成孔径雷达(SAR)穿透云层的结构信息。为解决跨模态域差异及现有视觉-语言模型在密集预测上的局限性,我们设计了两个核心模块:跨模态统一过程以实现多传感器表征对齐,以及双编码器融合模块,整合多个视觉基础模型的层次特征,实现文本对齐的多模态分割。大量实验表明,MM-OVSeg在不同云况下均展现出优异的鲁棒性与泛化性能。源数据集与代码已公开于 https://github.com/Jimmyxichen/MM-OVSeg。

原文摘要 · Abstract (English)

Open-vocabulary segmentation enables pixel-level recognition from an open set of textual categories, allowing generalization beyond fixed classes. Despite great potential in remote sensing, progress in this area remains largely limited to clear-sky optical data and struggles under cloudy or haze-contaminated conditions. We present MM-OVSeg, a multimodal Optical-SAR fusion framework for resilient open-vocabulary segmentation under adverse weather conditions. MM-OVSeg leverages the complementary strengths of the two modalities--optical imagery provides rich spectral semantics, while synthetic aperture radar (SAR) offers cloud-penetrating structural cues. To address the cross-modal domain gap and the limited dense prediction capability of current vision-language models, we propose two key designs: a cross-modal unification process for multi-sensor representation alignment, and a dual-encoder fusion module that integrates hierarchical features from multiple vision foundation models for text-aligned multimodal segmentation. Extensive experiments demonstrate that MM-OVSeg achieves superior robustness and generalization across diverse cloud conditions. The source dataset and code are available at https://github.com/Jimmyxichen/MM-OVSeg.

遥感分割多模态融合开放词汇SAR图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。