arXiv:2412.20742cs.CV2024-12被引 10

首个统一多时相遥感任务的视觉语言模型,支持单图、双时相图对和视频输入。

UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models

论文配图:UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models
图 1 · 摘自论文原文
  • 统一视觉表征框架,兼容单图、双时相图对与视频输入
  • 在遥感问答、变化描述生成等任务上达顶尖性能
  • 适合需要多时相遥感分析的研究者与应用开发者

遥感图像与自然图像之间的领域差异近年来受到广泛关注,视觉语言模型(VLMs)在遥感多模态任务中展现出优异的泛化能力。然而,现有研究仍缺乏对遥感VLM如何处理不同类型视觉输入的深入探索。为此,我们提出首个统一多时相遥感任务的视觉语言模型UniRS,支持单张图像、双时相图像对及视频作为输入,实现统一框架下的全方位遥感时序分析。采用统一视觉表征,使模型可接收多种视觉输入;针对双时相图像对任务,定制变化提取模块以增强时空特征捕捉能力。同时设计适配模型推理过程的提示增强机制,利用通用VLM的先验知识为UniRS提供推理线索。通过混合数据集联合微调促进多任务知识共享。实验表明,UniRS在视觉问答、变化描述生成、视频场景分类等多样化任务上均达到当前最优性能,充分验证其在统一多时相遥感任务中的泛化性与有效性。代码与数据集将不久后发布。

原文摘要 · Abstract (English)

The domain gap between remote sensing imagery and natural images has recently received widespread attention and Vision-Language Models (VLMs) have demonstrated excellent generalization performance in remote sensing multimodal tasks. However, current research is still limited in exploring how remote sensing VLMs handle different types of visual inputs. To bridge this gap, we introduce \textbf{UniRS}, the first vision-language model \textbf{uni}fying multi-temporal \textbf{r}emote \textbf{s}ensing tasks across various types of visual input. UniRS supports single images, dual-time image pairs, and videos as input, enabling comprehensive remote sensing temporal analysis within a unified framework. We adopt a unified visual representation approach, enabling the model to accept various visual inputs. For dual-time image pair tasks, we customize a change extraction module to further enhance the extraction of spatiotemporal features. Additionally, we design a prompt augmentation mechanism tailored to the model's reasoning process, utilizing the prior knowledge of the general-purpose VLM to provide clues for UniRS. To promote multi-task knowledge sharing, the model is jointly fine-tuned on a mixed dataset. Experimental results show that UniRS achieves state-of-the-art performance across diverse tasks, including visual question answering, change captioning, and video scene classification, highlighting its versatility and effectiveness in unifying these multi-temporal remote sensing tasks. Our code and dataset will be released soon.

遥感分析视觉语言模型多时相统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。