构建面向遥感的多模态大模型,解决复杂指令与像素级任务难题
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
- 用点集表示法统一遥感视觉语言数据,支持复杂条件理解
- 提出条件解析器与循环指代自增强策略,提升模型泛化能力
- 覆盖从分类到分割的多层级任务,适合遥感与地理科学研究者
视觉语言模型在遥感领域的应用已在场景分类、目标检测和图像描述等传统任务中展现出巨大潜力。然而,现有模型在处理复杂指令(如多重条件)或像素级操作(如分割与变化检测)时表现不佳。本文系统梳理了遥感视觉语言任务的层级认知需求,提出了遥感视觉语言任务集(RSVLTS),包含开放词汇任务(OVT)、指代表达任务(RET)、描述对象任务(DOT)及独立的视觉问答(VQA)。为应对挑战,提出基于点集的统一数据表示方法,结合条件解析器与基于循环指代的自增强策略,构建了GeoRSMLLM模型。该模型可有效处理RSVLTS中的多样化任务,推动遥感领域通用视觉语言解决方案的发展。
原文摘要 · Abstract (English)
The application of Vision-Language Models (VLMs) in remote sensing (RS) has demonstrated significant potential in traditional tasks such as scene classification, object detection, and image captioning. However, current models, which excel in Referring Expression Comprehension (REC), struggle with tasks involving complex instructions (e.g., exists multiple conditions) or pixel-level operations like segmentation and change detection. In this white paper, we provide a comprehensive hierarchical summary of vision-language tasks in RS, categorized by the varying levels of cognitive capability required. We introduce the Remote Sensing Vision-Language Task Set (RSVLTS), which includes Open-Vocabulary Tasks (OVT), Referring Expression Tasks (RET), and Described Object Tasks (DOT) with increased difficulty, and Visual Question Answering (VQA) aloneside. Moreover, we propose a novel unified data representation using a set-of-points approach for RSVLTS, along with a condition parser and a self-augmentation strategy based on cyclic referring. These features are integrated into the GeoRSMLLM model, and this enhanced model is designed to handle a broad range of tasks of RSVLTS, paving the way for a more generalized solution for vision-language tasks in geoscience and remote sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。