构建统一视觉语言模型,实现遥感多任务高效协同
Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- 设计动态分辨率与分块训练机制,适配遥感图像尺度差异
- 在多任务上超越现有遥感视觉语言模型,接近专用模型表现
- 开源全套工具与数据,支持遥感通用模型研究
随着Transformer在单个遥感任务中表现优异,构建统一的多任务学习(MTL)模型已成为可能。相比单任务方法,多任务学习具备更强泛化性、可扩展性和实用性。近年来,视觉语言模型(VLMs)在遥感图像理解、视觉定位和超高清(UHR)图像推理方面取得显著进展,文本统一接口展现出巨大潜力。本文提出RSCoVLM,一个简单而灵活的遥感多任务视觉语言模型基线。首先,构建数据工程体系,涵盖数据获取、离线处理与整合,以及在线加载与加权,有效应对复杂遥感数据环境,生成灵活的视觉-语言对话。其次,提出统一动态分辨率策略,解决遥感图像尺度多样性问题;针对超高清图像,引入Zoom-in Chain机制及配套数据集LRS-VQA-Zoom,减轻计算负担。此外,显著提升目标检测能力,并提出新评估协议,确保与传统检测模型公平比较。大量实验表明,RSCoVLM在多个任务上达到当前最优性能,优于现有遥感VLM,甚至媲美专业专家模型。所有训练与评估工具、模型权重及数据集均已开源,以支持可复现性。我们期望该基线推动通用遥感模型的发展。
原文摘要 · Abstract (English)
With Transformers achieving outstanding performance on individual remote sensing (RS) tasks, we are now approaching the realization of a unified model that excels across multiple tasks through multi-task learning (MTL). Compared to single-task approaches, MTL methods offer improved generalization, enhanced scalability, and greater practical applicability. Recently, vision language models (VLMs) have achieved promising results in RS image understanding, grounding, and ultra-high-resolution (UHR) image reasoning, respectively. Moreover, the unified text-based interface demonstrates significant potential for MTL. Hence, in this work, we present RSCoVLM, a simple yet flexible VLM baseline for RS MTL. Firstly, we create the data curation engine, including data acquisition, offline processing and integrating, as well as online loading and weighting. This data engine effectively addresses complex RS data enviroment and generates flexible vision-language conversations. Furthermore, we propose a unified dynamic-resolution strategy to address the diverse image scales inherent in RS imagery. For UHR images, we introduce the Zoom-in Chain mechanism together with its corresponding dataset, LRS-VQA-Zoom. The strategies are flexible and effectively mitigate the computational burdens. Additionally, we significantly enhance the model's object detection capability and propose a novel evaluation protocol that ensures fair comparison between VLMs and conventional detection models. Extensive experiments demonstrate that RSCoVLM achieves state-of-the-art performance across diverse tasks, outperforming existing RS VLMs and even rivaling specialized expert models. All the training and evaluating tools, model weights, and datasets have been fully open-sourced to support reproducibility. We expect that this baseline will promote further progress toward general-purpose RS models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。