用多模态大模型实现高分辨率卫星图像的农田分割,无需额外视觉解码器。
MAgSeg: Segmentation of Agricultural Landscapes in High-Resolution Satellite Imagery using Multimodal Large Language Models

- 采用无解码器架构,让大模型直接生成局部区域文本标记进行分割。
- 在三个全球南方国家数据集上,性能显著超越现有先进模型。
- 适合需要低成本标注、小规模农户农田制图的研究与应用。
全球南方地区的农业景观分割面临地块碎片化、类内差异大及标注数据稀缺等挑战。近年来,多模态大语言模型(MLLMs)在分割任务中取得进展,但现有方法存在上下文长度瓶颈和卫星特征理解的领域适配差距。为此,本文提出MAgSeg——一种新型的无解码器MLLM分割方法。该方法在架构上高效,使标准MLLM能直接对高分辨率卫星影像中的复杂小农户农业景观进行分割,无需辅助视觉解码器。我们设计了一种新的指令微调数据格式,支持在高分辨率卫星影像上进行可扩展的微调与后训练,使MAgSeg能在生成文本标记时利用全局图像上下文,仅针对图像中一个局部块输出结果。在覆盖三个全球南方国家的数据集上的大量评估表明,MAgSeg显著优于当前最先进的MLLM基线模型,为小农户农业环境制图提供了一种可扩展解决方案。
原文摘要 · Abstract (English)
Agricultural landscape segmentation in the Global South is challenging as it is characterized by fragmented plots, high intra-class variance, and a scarcity of labeled training data. Recent advances in segmentation have been made by Multimodal Large Language Models (MLLMs). However, current approaches encounter critical context length bottlenecks and a domain alignment gap in understanding satellite features. We address these limitations through MAgSeg, a novel, decoder-free MLLM segmentation approach. MAgSeg is an architecturally efficient approach that enables standard MLLMs to perform segmentation of complex smallholder agricultural landscapes from high-resolution satellite imagery, without requiring auxiliary vision decoders. We introduce a novel instruction tuning data format designed to enable scalable fine-tuning and post-training on high resolution satellite imagery, which enables MAgSeg to learn from the global context of the image while generating text tokens for only a patch within the image. Extensive evaluations on datasets spanning three countries in the Global South demonstrate that MAgSeg significantly outperforms state-of-the-art MLLM baselines, offering a scalable solution to map smallholder agricultural environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。