arXiv:2605.10769cs.CVcs.AI2026-05中稿 · CVPR

用多专家提示生成高质量遥感描述,提升图像分割精度。

MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation

论文配图:MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation
图 1 · 摘自论文原文
  • 用多个提示让大模型从不同视角生成遥感描述。
  • 动态融合文本语义,在三个公开数据集上表现最优。
  • 适合遥感图像语义分割与多模态模型研究者。

多模态融合图像与场景描述已在多个领域广泛应用。然而,在复杂遥感(RS)场景中,现有研究主要集中于将文本语义信息与视觉特征融合的架构优化,而忽视了高质量遥感描述的生成及其在多模态语义融合中的有效性。为此,我们提出动态多模态大模型混合专家感知引导的遥感场景分割方法(MPerS)。设计多种提示使多模态大模型(MLLMs,包括LLaVA、ChatGPT、Qwen)生成高质量遥感描述,实现从不同专家视角感知遥感场景。采用DINOv3提取地物的密集视觉表征。设计动态混合专家模块,自适应融合最有效的文本语义。构建语言查询引导注意力机制,利用文本语义引导视觉特征实现精确分割。该方法在三个公开遥感语义分割数据集上均取得优异性能。

原文摘要 · Abstract (English)

The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex remote sensing (RS) scenes, existing studies have predominantly concentrated on architectural optimizations for integrating textual semantic information with visual features, while largely neglecting the generation of high-quality RS captions and the investigation of their effectiveness in multimodal semantic fusion.In this context, we propose the Dynamic MLLM Mixture-of-Experts Perception-Guided Remote Sensing Scene Segmentation, referred to as MPerS.We design multiple prompts for MLLMs to generate high-quality RS captions, enabling MLLMs to perceive RS scenes from diverse expert perspectives. DINOv3 is employed to extract dense visual representations of land-covers.We design a Dynamic MixExperts module that adaptively integrates the most effective textual semantics. Linguistic Query Guided Attention is constructed to utilize textual semantic information to guide visual features for precise segmentation. The MLLMs include LLaVA, ChatGPT, and Qwen. Our method achieves superior performance on three public semantic segmentation RS datasets.

遥感分割多模态融合大模型应用动态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。