arXiv:2412.02531cs.CV2024-12被引 5

用大模型生成文本提升遥感图像分类,无需人工标注

Multimodal Remote Sensing Scene Classification Using VLMs and Dual-Cross Attention Networks

  • 用视觉语言模型自动生成文本描述作为辅助信息
  • 在五个数据集上均超越基线模型,零样本分类也有效
  • 适合遥感、地理信息等需要多模态融合的场景

遥感场景分类(RSSC)在土地利用与资源管理中有重要应用。传统基于图像的单模态方法常受限于类内差异大、类间相似度高。引入文本信息可增强语义理解,但人工标注成本高。本文提出一种新框架,利用大视觉语言模型(VLMs)自动生成文本描述作为辅助模态,避免昂贵的人工标注。为充分挖掘视觉与文本的互补性,设计双交叉注意力网络实现模态融合。在五个RSSC数据集上的定量与定性实验表明,该框架持续优于基线模型。验证了VLM生成文本相比人工标注的有效性,并在零样本分类场景中证明所学多模态表示可有效用于未见类别。本研究为遥感场景分类中利用文本信息开辟新路径,提供有前景的多模态融合结构,对后续研究具启发意义。代码已公开。

原文摘要 · Abstract (English)

Remote sensing scene classification (RSSC) is a critical task with diverse applications in land use and resource management. While unimodal image-based approaches show promise, they often struggle with limitations such as high intra-class variance and inter-class similarity. Incorporating textual information can enhance classification by providing additional context and semantic understanding, but manual text annotation is labor-intensive and costly. In this work, we propose a novel RSSC framework that integrates text descriptions generated by large vision-language models (VLMs) as an auxiliary modality without incurring expensive manual annotation costs. To fully leverage the latent complementarities between visual and textual data, we propose a dual cross-attention-based network to fuse these modalities into a unified representation. Extensive experiments with both quantitative and qualitative evaluation across five RSSC datasets demonstrate that our framework consistently outperforms baseline models. We also verify the effectiveness of VLM-generated text descriptions compared to human-annotated descriptions. Additionally, we design a zero-shot classification scenario to show that the learned multimodal representation can be effectively utilized for unseen class classification. This research opens new opportunities for leveraging textual information in RSSC tasks and provides a promising multimodal fusion structure, offering insights and inspiration for future studies. Code is available at: https://github.com/CJR7/MultiAtt-RSSC

遥感分类多模态融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。