arXiv:2604.00820cs.CV2026-04被引 1

为遥感视觉语言模型设计持续学习基准,解决新任务和模态下的遗忘问题。

Continual Vision-Language Learning for Remote Sensing: Benchmarking and Analysis

  • 构建包含207k图文对的持续学习基准CLeaRS,覆盖多种任务与传感器。
  • 实测发现现有模型在新增任务时出现严重遗忘,性能显著下降。
  • 适合关注遥感多模态持续学习、模型鲁棒性研究的学者使用。

当前遥感视觉语言模型(RS VLMs)在图像理解任务中表现优异,但依赖静态训练数据,难以适应不断涌现的新传感模态和下游任务,暴露出持续学习能力不足的核心挑战:如何在不发生灾难性遗忘的前提下持续适应。尽管该问题具有重要实际意义,但针对RS VLMs的持续学习研究仍处于空白,且缺乏专用基准。本文提出CLeaRS,一个面向遥感领域持续视觉语言学习的综合性基准。该基准包含10个精心筛选的数据子集,涵盖超过207,000张图像-文本对,覆盖多样化的解读任务、传感模态与应用场景。我们定义了三种评估协议:长时序、模态增量与任务增量设置,系统评估模型的持续适应能力。对多种视觉语言模型的广泛测试表明,所有设置下均存在灾难性遗忘。此外,现有持续学习方法在适配至RS VLMs后,在应对任务、指令及模态变化时效果有限。研究结果凸显出亟需开发专为RS VLMs设计的持续学习方法。

原文摘要 · Abstract (English)

Current remote sensing vision-language models (RS VLMs) demonstrate impressive performance in image interpretation but rely on static training data, limiting their ability to accommodate continuously emerging sensing modalities and downstream tasks. This exposes a fundamental challenge: enabling RS VLMs to continually adapt without catastrophic forgetting. Despite its practical importance, the continual learning capability of RS VLMs remains underexplored, and no dedicated benchmark currently exists. In this work, we present CLeaRS, a comprehensive benchmark for continual vision-language learning in remote sensing. CLeaRS comprises 10 curated subsets with over 207k image-text pairs, spanning diverse interpretation tasks, sensing modalities, and application scenarios. We further define three evaluation protocols: long-horizon, modality-incremental, and task-incremental settings, to systematically assess continual adaptation. Extensive benchmarking of diverse vision-language models reveals catastrophic forgetting across all settings. Moreover, representative continual learning methods, when adapted to RS VLMs, exhibit limited effectiveness in handling task, instruction, and modality transitions. Our findings underscore the need for developing continual learning methods tailored to RS VLMs.

遥感持续学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。