轻量适配器让遥感视觉问答更高效,三类模型对比揭示最优架构。
A Unified Framework for Efficient Remote Sensing Visual Question Answering: Adapting Dual, Hybrid, and Encoder-Decoder Architectures

- 在冻结主干网络中注入轻量瓶颈适配器,仅用5%参数实现快速适配。
- 在高分辨率RSVQA-x数据集上,混合架构FLAVA表现最佳,推理与检索能力均衡。
- 适合灾情评估、城市监测等资源受限的遥感应用,推动高效多模态分析。
遥感视觉问答(RSVQA)因航空影像分辨率高、多尺度目标分布复杂及语义丰富而面临独特挑战。尽管通用领域基础模型已取得显著成果,但其直接应用于RSVQA受制于巨大领域差异及全量微调带来的计算开销。本文对三种不同视觉语言模型架构——双编码器CLIP、编码器-解码器BLIP和混合架构FLAVA——进行了对比分析,提出统一的架构改造流程,将轻量级瓶颈适配器嵌入冻结主干网络的注意力与MLP层,仅使用不到5%的可训练参数即可实现快速适应。在高分辨率RSVQA-x数据集上的实验表明,所有适配模型均能收敛,其中混合架构FLAVA在多模态推理与信息检索之间表现出更优平衡性。研究结果为灾害评估与城市监测中的资源高效视觉问答建立了新基准。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) in the Remote Sensing (RS) domain presents unique challenges due to the high resolution, multi scale object distribution, and semantic complexity of aerial imagery. While general domain Foundation Models have achieved remarkable success, their direct application to RSVQA is hindered by massive domain shifts and the computationally prohibitive nature of full fine tuning. This study presents a comparative analysis of RS Adapter, a Parameter Efficient Fine Tuning (PEFT) strategy, applied across three distinct Vision Language Model (VLM) architectures: the Dual Encoder CLIP, the Encoder Decoder BLIP, and the Hybrid FLAVA. We introduce a unified architectural surgery pipeline that injects lightweight bottleneck adapters into the attention and MLP layers of frozen backbones, enabling rapid adaptation with less than 5 percent of trainable parameters. Experimental results on the high resolution RSVQA x dataset demonstrate that while all adapted models achieve convergence, the Hybrid FLAVA architecture offers a superior balance of multimodal reasoning and retrieval capabilities compared to its unimodal counterparts. Our findings establish a new baseline for resource efficient VQA in disaster assessment and urban monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。