arXiv:2412.05679cs.CV2024-12被引 37

提出统一遥感视觉语言模型,支持多粒度理解与多图分析

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

  • 采用面向粒度的专家混合架构,在不增加参数量前提下实现多尺度理解
  • 在10个遥感任务上达到领先性能,像素级分割与变化检测表现优异
  • 适用于遥感图像解析、变化监测等场景,尤其适合需要细粒度理解的研究者

遥感视觉语言模型在遥感图像理解任务中已取得显著进展,但在像素级理解及多图输入处理方面仍存在不足。本文提出RSUniVLM,一种统一的端到端遥感视觉语言模型,可实现图像级、区域级和像素级的全面视觉理解,并有效处理多图像分析任务,如变化检测与变化描述。为在不增加模型规模的前提下提升多粒度表征能力,设计了粒度导向的专家混合架构,模型参数控制在约10亿。构建了一个大规模遥感指令跟随数据集,涵盖目标定位、视觉问答、语义分割等多种任务,基于多个现有遥感与通用领域数据集。大量实验验证了该模型在各类遥感任务上的卓越表现,达到当前最优水平。代码与模型将开源。

原文摘要 · Abstract (English)

Remote Sensing Vision-Language Models (RS VLMs) have made much progress in the tasks of remote sensing (RS) image comprehension. While performing well in multi-modal reasoning and multi-turn conversations, the existing models lack pixel-level understanding and struggle with multi-image inputs. In this work, we propose RSUniVLM, a unified, end-to-end RS VLM designed for comprehensive vision understanding across multiple granularity, including image-level, region-level, and pixel-level tasks. RSUniVLM also performs effectively in multi-image analysis, with instances of change detection and change captioning. To enhance the model's ability to capture visual information at different levels without increasing model size, we design a novel architecture called Granularity-oriented Mixture of Experts to constraint the model to about 1 billion parameters. We also construct a large-scale RS instruction-following dataset based on a variety of existing datasets in both RS and general domain, encompassing various tasks such as object localization, visual question answering, and semantic segmentation. Substantial experiments have been conducted to validate the superiority of the proposed RSUniVLM up to state-of-the-art across various RS tasks. Code and model will be available at \href{https://github.com/xuliu-cyber/RSUniVLM}{here}.

遥感视觉语言模型多粒度理解专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。