arXiv:2505.23010cs.CV2025-05被引 11

用视觉语言模型提升遥感图像超分辨率,让重建更符合真实场景。

SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model

  • 引入预训练视觉语言模型提取图像语义信息
  • 在UCMerced数据集上达29.30dB PSNR,优于现有方法
  • 适合需要高精度遥感重建的科研与应用人员

高分辨率遥感影像在城市规划、环境监测等应用中至关重要,但受传感器和传输限制,实际获取的图像常存在分辨率退化问题。遥感图像超分辨率(RSISR)旨在从低分辨率(LR)输入重建高分辨率(HR)图像,是一种经济高效的替代方案。现有方法主要关注像素级低层特征,忽视了遥感场景的高层语义理解,可能导致重建结果语义不一致。为此,本文提出SeG-SR框架,利用视觉语言模型(VLMs)提取输入图像的语义知识,并用于引导超分辨率过程。具体包括:设计语义特征提取模块(SFEM),通过预训练VLM提取遥感图像语义;提出语义定位模块(SLM),生成一系列语义引导;构建可学习调制模块(LMM),将语义引导作用于超分网络特征,融入高层场景理解。在三个数据集上的实验表明,SeG-SR达到当前最优性能,且对多种超分架构均有稳定提升。特别地,在UCMerced数据集的x4超分任务中,取得29.3042 dB的PSNR和0.7961的SSIM。

原文摘要 · Abstract (English)

High-resolution (HR) remote sensing imagery plays a vital role in a wide range of applications, including urban planning and environmental monitoring. However, due to limitations in sensors and data transmission links, the images acquired in practice often suffer from resolution degradation. Remote Sensing Image Super-Resolution (RSISR) aims to reconstruct HR images from low-resolution (LR) inputs, providing a cost-effective and efficient alternative to direct HR image acquisition. Existing RSISR methods primarily focus on low-level characteristics in pixel space, while neglecting the high-level understanding of remote sensing scenes. This may lead to semantically inconsistent artifacts in the reconstructed results. Motivated by this observation, our work aims to explore the role of high-level semantic knowledge in improving RSISR performance. We propose a Semantic-Guided Super-Resolution framework, SeG-SR, which leverages Vision-Language Models (VLMs) to extract semantic knowledge from input images and uses it to guide the super resolution (SR) process. Specifically, we first design a Semantic Feature Extraction Module (SFEM) that utilizes a pretrained VLM to extract semantic knowledge from remote sensing images. Next, we propose a Semantic Localization Module (SLM), which derives a series of semantic guidance from the extracted semantic knowledge. Finally, we develop a Learnable Modulation Module (LMM) that uses semantic guidance to modulate the features extracted by the SR network, effectively incorporating high-level scene understanding into the SR pipeline. We validate the effectiveness and generalizability of SeG-SR through extensive experiments: SeG-SR achieves state-of-the-art performance on three datasets, and consistently improves performance across various SR architectures. Notably, for the x4 SR task on UCMerced dataset, it attained a PSNR of 29.3042 dB and an SSIM of 0.7961.

遥感图像超分辨率视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。