arXiv:2510.13993cs.CVcs.AI2025-10被引 2

融合视觉与视觉语言模型,提升遥感图像少样本学习的检测与理解能力。

Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models

  • 用YOLO结合LLaVA等VLM,实现遥感图像的上下文感知分析。
  • 飞机检测与计数平均MAE降低48.46%,在退化图像中仍表现优异。
  • 适合遥感领域少样本场景,尤其在标注数据稀缺时效果显著。

遥感在城市规划、环境监测和灾害响应等领域日益重要。尽管数据量激增,传统视觉模型受限于对大量特定领域标注数据的需求及其对复杂环境上下文的理解能力不足。视觉语言模型(VLM)通过融合视觉与文本信息提供了互补路径,但其在遥感中的应用仍不充分,尤其因具备通用性。本文研究将传统视觉模型与VLM(如LLaVA、ChatGPT、Gemini)结合,以增强遥感图像分析,聚焦飞机检测与场景理解。在标注与未标注遥感数据及退化图像场景下评估性能。结果表明,各类模型在飞机检测与计数上的平均MAE降低48.46%,尤其在原始与退化场景中表现突出;同时在综合理解方面,CLIPScore提升6.17%。该方法为少样本遥感图像分析提供了更高效、先进的解决方案。

原文摘要 · Abstract (English)

Remote sensing has become a vital tool across sectors such as urban planning, environmental monitoring, and disaster response. While the volume of data generated has increased significantly, traditional vision models are often constrained by the requirement for extensive domain-specific labelled data and their limited ability to understand the context within complex environments. Vision Language Models offer a complementary approach by integrating visual and textual data; however, their application to remote sensing remains underexplored, particularly given their generalist nature. This work investigates the combination of vision models and VLMs to enhance image analysis in remote sensing, with a focus on aircraft detection and scene understanding. The integration of YOLO with VLMs such as LLaVA, ChatGPT, and Gemini aims to achieve more accurate and contextually aware image interpretation. Performance is evaluated on both labelled and unlabelled remote sensing data, as well as degraded image scenarios which are crucial for remote sensing. The findings show an average MAE improvement of 48.46% across models in the accuracy of aircraft detection and counting, especially in challenging conditions, in both raw and degraded scenarios. A 6.17% improvement in CLIPScore for comprehensive understanding of remote sensing images is obtained. The proposed approach combining traditional vision models and VLMs paves the way for more advanced and efficient remote sensing image analysis, especially in few-shot learning scenarios.

遥感少样本学习视觉语言模型目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。