用大模型直接看卫星图答问题,更准更流畅。
Large Vision-Language Models for Remote Sensing Visual Question Answering
- 用两阶段训练让大模型同时理解图像和问题
- 在RSVQAxBEN数据集上超越现有方法
- 适合需要自然语言回答的遥感分析场景
遥感视觉问答(RSVQA)是一项挑战性任务,需从复杂卫星图像中解读信息并回答自然语言问题。传统方法依赖独立的视觉特征提取器和语言处理模型,计算成本高且难以应对开放式问题。本文提出一种新方法,利用生成式大视觉语言模型(LVLM)简化RSVQA流程。该方法采用两阶段训练策略:领域自适应预训练与基于提示的微调,使模型能根据图像和文本输入生成自然语言答案,无需预定义答案类别。我们在RSVQAxBEN数据集上评估模型,性能优于现有最先进基线。此外,人工评估显示,我们的方法生成的答案更准确、相关且流畅。结果表明生成式LVLM在遥感分析领域具有巨大潜力。
原文摘要 · Abstract (English)
Remote Sensing Visual Question Answering (RSVQA) is a challenging task that involves interpreting complex satellite imagery to answer natural language questions. Traditional approaches often rely on separate visual feature extractors and language processing models, which can be computationally intensive and limited in their ability to handle open-ended questions. In this paper, we propose a novel method that leverages a generative Large Vision-Language Model (LVLM) to streamline the RSVQA process. Our approach consists of a two-step training strategy: domain-adaptive pretraining and prompt-based finetuning. This method enables the LVLM to generate natural language answers by conditioning on both visual and textual inputs, without the need for predefined answer categories. We evaluate our model on the RSVQAxBEN dataset, demonstrating superior performance compared to state-of-the-art baselines. Additionally, a human evaluation study shows that our method produces answers that are more accurate, relevant, and fluent. The results highlight the potential of generative LVLMs in advancing the field of remote sensing analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。