arXiv:2512.24022cs.CVcs.AI2025-12被引 3

融合多尺度视觉特征,提升遥感图像理解能力

FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing

  • 设计多特征融合机制,结合全局与局部视觉信息
  • 在多个遥感任务上达到领先或接近顶尖水平
  • 适合遥感图像分析、智能解译等研究者使用

大型视觉语言模型在多种任务中表现优异,但在遥感领域面临挑战,因遥感图像与自然图像存在本质差异。现有遥感视觉语言模型难以提取细粒度视觉特征,且在深度语言处理中易发生视觉遗忘。为此,我们提出MF-RSVLM,一种多特征融合遥感视觉语言模型,能有效提取并融合视觉特征以增强遥感理解。该模型学习多尺度视觉表征,结合全局上下文与局部细节,更好捕捉遥感场景中的小尺度与复杂结构。通过循环视觉特征注入机制,确保语言模型始终基于视觉证据,减少生成过程中的视觉遗忘。在多个遥感基准测试上的大量实验表明,MF-RSVLM在遥感分类、图像描述和视觉问答任务中均达到当前最优或具有竞争力的性能。代码已公开于https://github.com/Yunkaidang/RSVLM。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain due to the inherent differences between remote sensing images and natural images. Existing remote sensing VLMs often fail to extract fine-grained visual features and suffer from visual forgetting during deep language processing. To address this, we introduce MF-RSVLM, a Multi-Feature Fusion Remote Sensing Vision--Language Model that effectively extracts and fuses visual features for RS understanding. MF-RSVLM learns multi-scale visual representations and combines global context with local details, improving the capture of small and complex structures in RS scenes. A recurrent visual feature injection scheme ensures the language model remains grounded in visual evidence and reduces visual forgetting during generation. Extensive experiments on diverse RS benchmarks show that MF-RSVLM achieves state-of-the-art or highly competitive performance across remote sensing classification, image captioning, and VQA tasks. Our code is publicly available at https://github.com/Yunkaidang/RSVLM.

遥感图像视觉语言模型多特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。