首个支持遥感图像像素级定位的多模态大模型,实现高精度视觉对话。
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
- 构建端到端遥感多模态模型,支持像素级视觉定位与对话生成
- 在4K高清分辨率下表现优异,单目标与多目标分割均超越现有模型
- 自建地理标注数据集GeoPixelD,支持遥感领域细粒度对话理解
大型多模态模型(LMMs)虽已重视细粒度定位对视觉理解的重要性,但其在遥感(RS)图像中的应用受限于视角、尺度变化及小目标等挑战。现有模型难以实现区域级精细理解,且缺乏细粒度、领域特定的标注数据制约了遥感对话能力的发展。为此,我们提出GeoPixel——首个支持像素级定位的高分辨率遥感多模态大模型,可在任意长宽比下处理高达4K HD分辨率图像。通过半自动化流程构建了面向遥感的视觉定位数据集GeoPixelD,采用标记集合提示与空间先验控制数据生成。GeoPixel在单目标和多目标分割任务中均显著优于现有LMMs,消融实验证明各模块有效性。代码与数据将公开发布。
原文摘要 · Abstract (English)
Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image domain, and these models perform poorly for remote sensing (RS). The distinct overhead viewpoint, scale variation, and presence of small objects in high-resolution RS imagery present a unique challenge in region-level comprehension. Moreover, the development of the grounding conversation capability of LMMs within RS is hindered by the lack of granular, RS domain-specific grounded data. Addressing these limitations, we propose GeoPixel - the first end-to-end high resolution RS-LMM that supports pixel-level grounding. This capability allows fine-grained visual perception by generating interleaved masks in conversation. GeoPixel supports up to 4K HD resolution in any aspect ratio, ideal for high-precision RS image analysis. To support the grounded conversation generation (GCG) in RS imagery, we curate a visually grounded dataset GeoPixelD through a semi-automated pipeline that utilizes set-of-marks prompting and spatial priors tailored for RS data to methodically control the data generation process. GeoPixel demonstrates superior performance in pixel-level comprehension, surpassing existing LMMs in both single-target and multi-target segmentation tasks. Our methodological ablation studies validate the effectiveness of each component in the overall architecture. Our code and data will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。