Aquila通过多尺度高分辨率特征融合,提升遥感图像理解能力。
Aquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image Comprehension
- 引入可学习的分层空间特征融合模块,支持高分辨率输入
- 在多层语言模型中重复集成,实现深层视觉语言对齐
- 兼顾遥感图像解析与自然语言处理性能,适合遥感领域应用
近期大规模视觉语言模型(VLMs)通过视觉指令微调在视觉语言理解方面取得显著进展,展现出在遥感图像解析中的巨大潜力。然而,现有遥感视觉语言模型(RSVLMs)往往难以捕捉遥感场景的复杂特性,因其通常依赖低分辨率、单尺度视觉特征,并采用简单方法将视觉特征映射到语言特征。本文提出Aquila,一种先进的视觉语言基础模型,旨在增强遥感图像的视觉特征表达并实现更精确的视觉-语言特征对齐。该方法引入可学习的分层空间特征融合(SFI)模块,支持高分辨率图像输入,聚合多尺度视觉特征,实现复杂视觉信息的细节表征。同时,SFI模块被反复集成至大型语言模型(LLM)各层,实现深度视觉-语言特征对齐,且不损害模型在自然语言任务中的性能。这些创新通过更高分辨率和多尺度输入捕捉更精细的视觉效果,并显著增强特征对齐,有效提升模型从图文数据中学习的能力。我们通过大量定量实验与定性分析验证了Aquila的有效性,其性能显著优于现有方法。
原文摘要 · Abstract (English)
Recently, large vision language models (VLMs) have made significant strides in visual language capabilities through visual instruction tuning, showing great promise in the field of remote sensing image interpretation. However, existing remote sensing vision language models (RSVLMs) often fall short in capturing the complex characteristics of remote sensing scenes, as they typically rely on low resolution, single scale visual features and simplistic methods to map visual features to language features. In this paper, we present Aquila, an advanced visual language foundation model designed to enable richer visual feature representation and more precise visual-language feature alignment for remote sensing images. Our approach introduces a learnable Hierarchical Spatial Feature Integration (SFI) module that supports high resolution image inputs and aggregates multi scale visual features, allowing for the detailed representation of complex visual information. Additionally, the SFI module is repeatedly integrated into the layers of the large language model (LLM) to achieve deep visual language feature alignment, without compromising the model's performance in natural language processing tasks. These innovations, capturing detailed visual effects through higher resolution and multi scale input, and enhancing feature alignment significantly improve the model's ability to learn from image text data. We validate the effectiveness of Aquila through extensive quantitative experiments and qualitative analyses, demonstrating its superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。