融合文本与视觉信息,实现遥感图像开放词汇语义分割
Open-Vocabulary Semantic Segmentation Network Integrating Object-Level Label and Scene-Level Semantic Features for Multimodal Remote Sensing Images
- 用双分支文本编码器提取场景与目标级语义
- 在多个数据集上达到领先精度,跨区域泛化能力强
- 适合需要可解释性的遥感分析场景
多模态遥感图像语义分割在土地利用/覆盖制图、环境监测和精准地球观测中至关重要。现有方法多关注视觉模态的融合,却忽视了非视觉文本数据这一丰富知识源,难以弥合视觉模式与真实概念间的语义鸿沟。为此,我们提出TSMNet,一种文本监督的多模态开放词汇语义分割网络,通过协同整合文本监督与视觉表示,实现开放词汇分割。不同于传统框架,TSMNet引入双分支文本编码器,从多种文本数据中提取场景级语义与目标级标签信息,实现动态跨模态融合。文本衍生特征通过提出的文本引导视觉语义融合模块与视觉嵌入动态交互,实现领域感知的特征优化与人类可解释决策。为验证方法,我们创新构建两个新多模态数据集,并与多个SOTA模型进行广泛对比。结果表明,TSMNet在多个地理与传感器场景下均取得优异分割精度,展现出强鲁棒性。该工作建立了可解释遥感分析的新范式,证明文本知识融合显著提升模型泛化能力。代码将开源于https://github.com/yeyuanxin110/TSMNet。
原文摘要 · Abstract (English)
Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. Current multi-modal approaches mainly focus on integrating complementary visual modalities, yet neglect the incorporating of non-visual textual data - a rich source of knowledge that can bridge semantic gaps between visual patterns and real-world concepts. To address this limitation, we propose TSMNet, a text supervised multi-modal open vocabulary semantic segmentation network that synergistically integrates textual supervision with visual representation for open-vocabulary semantic segmentation. Unlike conventional multi-modal segmentation frameworks, TSMNet introduces a dual-branch text encoder to extract both scene-level semantic and object-level label information from various textual data, enabling dynamic cross-modal fusion. These text-derived features dynamically interact with visual embeddings through the proposed text-guided visual semantic fusion module, enabling domain-aware feature refinement and human-interpretable decision-making. To verify our method, we innovatively construct two new multi-modal datasets, and carry out extensive experiments to make a comprehensive comparison between the proposed method and other state-of-the-art (SOTA) semantic segmentation models. Results demonstrate that TSMNet achieves superior segmentation accuracy while exhibiting robust generalization capabilities across diverse geographical and sensor-specific scenarios. This work establishes a new paradigm for explainable remote sensing analysis, demonstrating that textual knowledge integration significantly enhances model generalizability. The source code will be available at https://github.com/yeyuanxin110/TSMNet
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。