arXiv:2508.11919cs.CV2025-08中稿 · ed被引 1

无需文字标注,用时间序列对齐遥感图像,提升中分辨率影像理解能力。

TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing

  • 通过跨视角时间对比学习,对齐哨兵2号时序影像与地面图像。
  • 在无文本标注下实现零样本分类,验证单像素时序信息有效性。
  • 专为中分辨率遥感设计,突出光谱与时间特征而非空间细节。

视觉-语言模型(VLM)在遥感应用中展现出巨大潜力,尤其在无需标注的零样本土地利用/土地覆盖(LULC)分类与检索任务中表现优异。然而,现有方法存在依赖基于描述的监督信号的问题,这类数据常缺失或语义覆盖不足;同时,它们多源自适用于高分辨率图像的通用架构,更关注空间上下文而忽略光谱与时间信息,难以适配中分辨率遥感影像。为此,本文提出TimeSenCLIP,一种面向遥感时序数据的轻量级视觉-语言模型,采用跨视角时间对比学习框架,将多光谱哨兵2号(Sentinel-2)时序影像与地理标记的地面图像进行对齐,无需任何文本标注。不同于以往模型,TimeSenCLIP强调时间与光谱信号,探究单像素时序序列是否足以支持多种任务。实验表明,该方法在无监督条件下有效捕捉时序动态特征,显著提升中分辨率影像的理解能力。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown significant promise in remote sensing applications, particularly for land-use and land-cover (LULC) mapping via zero-shot classification and retrieval. However, current approaches face several key challenges, such as the dependence on caption-based supervision, which is often not available or very limited in terms of the covered semantics, and the fact of being adapted from generic VLM architectures that are suitable for very high resolution images. Consequently, these models tend to prioritize spatial context over spectral and temporal information, limiting their effectiveness for medium-resolution remote sensing imagery. In this work, we present TimeSenCLIP, a lightweight VLM for remote sensing time series, using a cross-view temporal contrastive framework to align multispectral Sentinel-2 time series with geo-tagged ground-level imagery, without requiring textual annotations. Unlike prior VLMs, TimeSenCLIP emphasizes temporal and spectral signals over spatial context, investigating whether single-pixel time series contain sufficient information for solving a variety of tasks.

遥感时序建模视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。