arXiv:2505.21375cs.CV2025-05NeurIPS被引 41

首个支持8K遥感图像的多模态大模型,解决高分辨率数据与计算瓶颈问题。

GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution

  • 通过筛选关键对象区域,减少冗余背景信息以降低计算开销。
  • 在8192×8192分辨率下仍保持优异性能,超越现有遥感多模态模型。
  • 适合遥感分析、地球观测等需要高精度视觉理解的研究者使用。

超高清遥感影像为地球观测提供宝贵数据,但现有多模态基础模型面临两大瓶颈:(1) 超高分辨率训练数据稀缺;(2) 图像尺寸过大导致的令牌爆炸。为缓解数据不足,我们构建了目前最高分辨率的遥感视觉-语言数据集SuperRS-VQA(平均8,376×8,376)和HighRS-VQA(平均2,000×1,912),涵盖22个真实对话任务。针对令牌爆炸问题,初步研究发现遥感图像中关键信息集中于少数目标中心令牌,而剪除背景令牌(如海洋或森林)反而能提升性能。据此提出背景令牌剪枝与锚定令牌选择策略,在降低内存占用的同时保留核心语义。融合上述方法,我们推出GeoLLaVA-8K,首个可处理8,192×8,192分辨率输入的遥感专用多模态大模型,基于LLaVA框架,训练于SuperRS-VQA与HighRS-VQA,在XLRS-Bench上达到新基准。

原文摘要 · Abstract (English)

Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To address data scarcity, we introduce SuperRS-VQA (avg. 8,376$\times$8,376) and HighRS-VQA (avg. 2,000$\times$1,912), the highest-resolution vision-language datasets in RS to date, covering 22 real-world dialogue tasks. To mitigate token explosion, our pilot studies reveal significant redundancy in RS images: crucial information is concentrated in a small subset of object-centric tokens, while pruning background tokens (e.g., ocean or forest) can even improve performance. Motivated by these findings, we propose two strategies: Background Token Pruning and Anchored Token Selection, to reduce the memory footprint while preserving key semantics.Integrating these techniques, we introduce GeoLLaVA-8K, the first RS-focused multimodal large language model capable of handling inputs up to 8K$\times$8K resolution, built on the LLaVA framework. Trained on SuperRS-VQA and HighRS-VQA, GeoLLaVA-8K sets a new state-of-the-art on the XLRS-Bench.

遥感多模态大模型8K分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。