用稀疏提示让大模型专注遥感图像关键区域,避免被背景干扰。
GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing
- 基于视觉块设计空间结构软提示,动态聚焦任务相关区域。
- 在多个遥感问答数据集上表现优于现有方法,参数效率高。
- 适合需要高效适配大模型的遥感图像分析场景。
近年来,多模态大语言模型(MLLMs)在视觉问答任务中取得显著进展。然而,直接将现有微调方法应用于遥感(RS)图像时,常因背景噪声过拟合或忽略目标细节而效果不佳。这主要源于遥感图像固有的大规模变化、目标分布稀疏及复杂区域语义特征。为应对这些挑战,我们提出一种参数高效微调策略——引导式区域感知稀疏提示(GRASP)。GRASP 在冻结的视觉标记网格中提取的空间块基础上,引入空间结构化的软提示,并通过问题引导的稀疏融合机制,动态聚合任务特定上下文形成紧凑全局提示,使模型能聚焦相关区域并过滤背景噪声。在多个遥感视觉问答基准上的大量实验表明,GRASP 在性能上可与现有微调和提示方法媲美,同时保持高参数效率。
原文摘要 · Abstract (English)
In recent years, Multimodal Large Language Models (MLLMs) have made significant progress in visual question answering tasks. However, directly applying existing fine-tuning methods to remote sensing (RS) images often leads to issues such as overfitting on background noise or neglecting target details. This is primarily due to the large-scale variations, sparse target distributions, and complex regional semantic features inherent in RS images. These challenges limit the effectiveness of MLLMs in RS tasks. To address these challenges, we propose a parameter-efficient fine-tuning (PEFT) strategy called Guided Region-Aware Sparse Prompting (GRASP). GRASP introduces spatially structured soft prompts associated with spatial blocks extracted from a frozen visual token grid. Through a question-guided sparse fusion mechanism, GRASP dynamically aggregates task-specific context into a compact global prompt, enabling the model to focus on relevant regions while filtering out background noise. Extensive experiments on multiple RSVQA benchmarks show that GRASP achieves competitive performance compared to existing fine-tuning and prompt-based methods while maintaining high parameter efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。