专为遥感图像设计的多模态大模型,提升视觉语言对齐与空间理解能力。
LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation
- 改进视觉编码器与桥接层,实现高效视觉压缩与更好对齐
- 构建大规模遥感图文数据集,提升模型对地观测理解能力
- 适合遥感分析、地理信息研究等需要精准空间理解的场景
自动快速理解地球表面是掌握生态环境与科学决策的基础,亟需一个具备全面能力的统一系统来应对多样化需求。多模态大语言模型(MLLM)在智能遥感观测中展现出巨大潜力,可支持类人对话、统一图像理解、遵循多样指令并提供深度反馈。本文提出LHRS-Bot-Nova,一种专用于遥感(RS)图像理解的MLLM,能精准执行多种符合人类指令的遥感理解任务。该模型采用增强型视觉编码器与新型桥接层,实现高效视觉压缩与更优的语言-视觉对齐。为强化遥感导向的对齐能力,我们提出基于特征引导的图像重描述生成的大规模遥感图文数据集,并构建专门提升空间识别能力的指令数据集。大量实验表明,LHRS-Bot-Nova在各类遥感图像理解任务中表现优异。我们还通过复杂多选题评估基准,对比多种MLLM在复杂遥感感知与指令跟随中的性能,为未来模型选择与优化提供可靠依据。数据、代码与模型将开源于https://github.com/NJU-LHRS/LHRS-Bot。
原文摘要 · Abstract (English)
Automatically and rapidly understanding Earth's surface is fundamental to our grasp of the living environment and informed decision-making. This underscores the need for a unified system with comprehensive capabilities in analyzing Earth's surface to address a wide range of human needs. The emergence of multimodal large language models (MLLMs) has great potential in boosting the efficiency and convenience of intelligent Earth observation. These models can engage in human-like conversations, serve as unified platforms for understanding images, follow diverse instructions, and provide insightful feedbacks. In this study, we introduce LHRS-Bot-Nova, an MLLM specialized in understanding remote sensing (RS) images, designed to expertly perform a wide range of RS understanding tasks aligned with human instructions. LHRS-Bot-Nova features an enhanced vision encoder and a novel bridge layer, enabling efficient visual compression and better language-vision alignment. To further enhance RS-oriented vision-language alignment, we propose a large-scale RS image-caption dataset, generated through feature-guided image recaptioning. Additionally, we introduce an instruction dataset specifically designed to improve spatial recognition abilities. Extensive experiments demonstrate superior performance of LHRS-Bot-Nova across various RS image understanding tasks. We also evaluate different MLLM performances in complex RS perception and instruction following using a complicated multi-choice question evaluation benchmark, providing a reliable guide for future model selection and improvement. Data, code, and models will be available at https://github.com/NJU-LHRS/LHRS-Bot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。