针对遥感图像设计新模型,让视觉与语言理解更精准
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
- 用检索增强语义信息,实现多层级视觉特征对齐
- 在遥感场景分类和问答任务中显著优于通用模型
- 适合遥感分析、地理信息等领域的研究者使用
大型视觉语言模型(LVLM)在自然图像的跨模态任务中表现优异,但在遥感(RS)领域应用仍有限,主要因视觉表征、物体尺度与语义存在显著差异。这些差异阻碍了对遥感场景中从粗到细多层次语义信息的理解。为此,我们提出一种专为遥感设计的新型LVLM框架,包含两个核心组件:语义增强的多层级对齐与语义感知专家建模。首先,通过基于检索的语义增强模块,从遥感语义知识库中获取相关语义线索,融合用户查询与多层级视觉特征,生成跨层级的语义丰富表示。其次,构建语义专家,分别处理不同层次的语义表示,实现从粗粒度到细粒度的分层理解。在多个遥感任务(包括场景分类、视觉问答等)上的评估表明,该框架在多语义层级上均取得一致提升,有效弥合通用模型与遥感特定需求之间的差距。
原文摘要 · Abstract (English)
Large Vision and Language Models (LVLMs) have shown strong performance across various vision-language tasks in natural image domains. However, their application to remote sensing (RS) remains underexplored due to significant domain differences in visual appearances, object scales, and semantics. These discrepancies hider the effective understanding of RS scenes, which contain rich, multi-level semantic information spanning from coarse-to-fine levels. Hence, it limits the direct adaptation of existing LVLMs to RS imagery. To address this gap, we propose a novel LVLM framework tailored for RS understanding, incorporating two core components: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling. First, to align multi-level visual features, we introduce the retrieval-based Semantic Augmentation Module which enriches the visual features with relevant semantics across fine-to-coarse levels (e.g., object- and scene-level information). It is designed to retrieve relevant semantic cues from a RS semantic knowledge database, followed by aggregation of semantic cues with user query and multi-level visual features, resulting in semantically enriched representation across multiple levels. Second, for Semantic-aware Expert Modeling, we design semantic experts, where each expert is responsible for processing semantic representation at different levels separately. This enables hierarchical semantic understanding from coarse to fine levels. Evaluations across multiple RS tasks-including scene classification and VQA, etc.-demonstrate that the proposed framework achieves consistent improvements across multiple semantic levels. This highlights its capability and effectiveness in bridging the gap between general LVLMs and unique demands of RS-specific vision-language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。