arXiv:2505.02829cs.AI2025-05NeurIPS被引 10

让卫星图像理解更聪明:能听懂复杂指令并精准分割目标

LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery

论文配图:LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery
图 1 · 摘自论文原文
  • 用语言指令引导,实现对遥感图像中多类目标的智能分割
  • 在描述任务上比RS-GPT4V高10.04%(BLEU-4),分割精度提升143.36%(gIoU)
  • 专为遥感设计的新数据集GRES,支持复杂语义推理

分割模型可识别图像中预定义的一组对象。然而,能够理解隐含涉及多个感兴趣对象的复杂用户查询的模型仍处于初级阶段。近期在推理分割方面的进展表明,视觉-语言模型可在开放域中运行并生成合理输出。但我们的实验显示,这类模型在复杂遥感图像上表现不佳。本文提出LISAt,一种专为描述复杂遥感场景、回答相关问题及分割目标而设计的视觉-语言模型。我们在新构建的地理空间推理-分割数据集GRES(含9,205张图像、27,615个标注)和包含超百万问答对的PreGRES多模态预训练数据集上训练LISAt。实验表明,LISAt在遥感描述任务上优于现有地理空间基础模型RS-GPT4V超过10.04%(BLEU-4),在推理分割任务上超越当前最先进开放域模型达143.36%(gIoU)。模型、数据集与代码已开源。

原文摘要 · Abstract (English)

Segmentation models can recognize a pre-defined set of objects in images. However, models that can reason over complex user queries that implicitly refer to multiple objects of interest are still in their infancy. Recent advances in reasoning segmentation--generating segmentation masks from complex, implicit query text--demonstrate that vision-language models can operate across an open domain and produce reasonable outputs. However, our experiments show that such models struggle with complex remote-sensing imagery. In this work, we introduce LISAt, a vision-language model designed to describe complex remote-sensing scenes, answer questions about them, and segment objects of interest. We trained LISAt on a new curated geospatial reasoning-segmentation dataset, GRES, with 27,615 annotations over 9,205 images, and a multimodal pretraining dataset, PreGRES, containing over 1 million question-answer pairs. LISAt outperforms existing geospatial foundation models such as RS-GPT4V by over 10.04 % (BLEU-4) on remote-sensing description tasks, and surpasses state-of-the-art open-domain models on reasoning segmentation tasks by 143.36 % (gIoU). Our model, datasets, and code are available at https://lisat-bair.github.io/LISAt/

遥感分割视觉语言模型指令理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。