SkyMoE用专家混合机制提升遥感多任务理解能力
SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts
- 引入自适应路由,按任务和粒度分配专用专家
- 在21个数据集上超越现有模型,实现多粒度精准理解
- 适合需要细粒度遥感分析的科研与应用人员
大型视觉语言模型(VLMs)显著提升了地理空间解释的效率与灵活性,但通用VLM在遥感(RS)任务中表现仍不理想。现有地理空间VLM普遍采用统一建模策略,难以区分任务类型与解释粒度,限制了局部细节感知与全局上下文理解的平衡。本文提出SkyMoE,一种面向多模态、多任务遥感解释的专家混合(MoE)视觉语言模型。其采用自适应路由器生成任务与粒度感知的路由指令,使特定大语言模型专家处理不同子任务。为增强专家解耦与粒度敏感性,引入上下文解耦增强策略,通过对比局部与全局特征对,引导专家进行层级化表征学习。同时构建MGRS-Bench基准,覆盖多种遥感解释任务与粒度级别,评估复杂场景下的泛化能力。在21个公开数据集上的大量实验表明,SkyMoE在各项任务中均达到领先性能,验证了其在遥感领域中的适应性、可扩展性与卓越的多粒度理解能力。
原文摘要 · Abstract (English)
The emergence of large vision-language models (VLMs) has significantly enhanced the efficiency and flexibility of geospatial interpretation. However, general-purpose VLMs remain suboptimal for remote sensing (RS) tasks. Existing geospatial VLMs typically adopt a unified modeling strategy and struggle to differentiate between task types and interpretation granularities, limiting their ability to balance local detail perception and global contextual understanding. In this paper, we present SkyMoE, a Mixture-of-Experts (MoE) vision-language model tailored for multimodal, multi-task RS interpretation. SkyMoE employs an adaptive router that generates task- and granularity-aware routing instructions, enabling specialized large language model experts to handle diverse sub-tasks. To further promote expert decoupling and granularity sensitivity, we introduce a context-disentangled augmentation strategy that creates contrastive pairs between local and global features, guiding experts toward level-specific representation learning. We also construct MGRS-Bench, a comprehensive benchmark covering multiple RS interpretation tasks and granularity levels, to evaluate generalization in complex scenarios. Extensive experiments on 21 public datasets demonstrate that SkyMoE achieves state-of-the-art performance across tasks, validating its adaptability, scalability, and superior multi-granularity understanding in remote sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。