arXiv:2501.06828cs.CV2025-01被引 13

让遥感图像理解突破像素级,支持对话式分割

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

  • 引入可学习记忆模块,实现跨实例的类级地理上下文捕捉
  • 构建6.5万张图14万实例的大规模标注数据集
  • 支持用户指令生成分割掩码,适合遥感智能分析场景

多模态大语言模型在遥感图像的图像级和区域级理解任务中取得显著进展,如图像描述、视觉问答和视觉定位。然而,现有遥感多模态模型缺乏像素级对话能力,即根据用户指令生成特定目标的分割掩码。本文提出GeoPix,一种将遥感图像理解扩展至像素级别的多模态大语言模型。通过在大语言模型中集成掩码预测器,将视觉编码器提取的特征转化为条件于语言模型分割标记嵌入的掩码。为应对遥感影像中多尺度目标的分割需求,掩码预测器融合了类别感知的可学习记忆模块,以在全数据集范围内捕获并存储实例级的类级地理上下文。此外,针对缺乏大规模训练数据的问题,构建了GeoPixInstruct数据集,包含65,463张图像和140,412个实例,每个实例配有文本描述、边界框和掩码。同时,设计两阶段训练策略,平衡文本生成与掩码预测在多任务优化中的不同需求。大量实验验证了GeoPix在像素级分割任务上的有效性与优越性,同时在图像级和区域级基准测试中保持竞争力。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visual question answering, and visual grounding. However, existing RS MLLMs lack the pixel-level dialogue capability, which involves responding to user instructions with segmentation masks for specific instances. In this paper, we propose GeoPix, a RS MLLM that extends image understanding capabilities to the pixel level. This is achieved by equipping the MLLM with a mask predictor, which transforms visual features from the vision encoder into masks conditioned on the LLM's segmentation token embeddings. To facilitate the segmentation of multi-scale objects in RS imagery, a class-wise learnable memory module is integrated into the mask predictor to capture and store class-wise geo-context at the instance level across the entire dataset. In addition, to address the absence of large-scale datasets for training pixel-level RS MLLMs, we construct the GeoPixInstruct dataset, comprising 65,463 images and 140,412 instances, with each instance annotated with text descriptions, bounding boxes, and masks. Furthermore, we develop a two-stage training strategy to balance the distinct requirements of text generation and masks prediction in multi-modal multi-task optimization. Extensive experiments verify the effectiveness and superiority of GeoPix in pixel-level segmentation tasks, while also maintaining competitive performance in image- and region-level benchmarks.

遥感图像像素分割多模态模型大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。