提出一种自适应的遥感图像令牌剪枝方法,提升大模型效率与精度。
SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models

- 根据任务需求动态调整分辨率,实现细粒度令牌剪枝。
- 在XLRS-Bench上比GeoLLaVA-8K高2.3%准确率,推理速度提升2.4倍。
- 适合需要高效处理高分辨率遥感图像的科研与应用人员。
遥感视觉语言模型(RS-LVLMs)虽提升了对地球观测图像的多模态理解能力,但其性能受限于高分辨率处理带来的计算开销——视觉令牌数量随输入分辨率线性增长而呈平方级上升,而关键视觉证据本身稀疏且在扩展序列中逐渐稀释。现有令牌剪枝方法多依赖无尺度感知的分辨率策略和孤立的重要性线索,难以实现任务对齐的粒度自适应与整体地理空间证据保留。为此,我们提出规模自适应、地理空间证据调制的令牌剪枝框架SA-GEM,可即插即用。该框架通过轻量级路由模块根据查询动态选择分辨率,同时利用令牌重要性调制器联合建模任务相关性、空间结构与局部冗余性,以保全整体地理空间证据。实验表明,更高分辨率并非普遍有益;一旦达到足够粒度,令牌质量远胜数量。在多个基准测试中,SA-GEM均优于现有剪枝方法。在XLRS-Bench上,其准确率超越GeoLLaVA-8K 2.3%,总推理速度提升2.4倍。
原文摘要 · Abstract (English)
RS-LVLMs have advanced multimodal understanding of Earth observation imagery, yet their performance is fundamentally constrained by high-resolution processing, as visual token counts grow quadratically with linear input resolution while important visual evidence is inherently sparse and increasingly diluted across the expanded sequence. Existing token pruning methods largely rely on scale-agnostic resolution policies and isolated importance cues, limiting task-aligned granularity adaptation and holistic evidence preservation. To address this, we present Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning (SA-GEM), a plug-and-play framework that unifies task-adaptive token granularity allocation with holistic geospatial token importance modulation. Specifically, a lightweight router selects the resolution based on query-dependent token granularity, while a token importance modulator jointly models task relevance, spatial structure, and local redundancy to preserve holistic geospatial evidence. We show that higher resolution is not universally beneficial and, once sufficient granularity is reached, token quality matters more than token quantity. Experiments across various benchmarks demonstrate that SA-GEM achieves consistent gains in both accuracy and efficiency over existing pruning methods. On XLRS-Bench, it surpasses GeoLLaVA-8K by 2.3% in accuracy with a 2.4 times total inference speedup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。