arXiv:2505.08723cs.CV2025-05被引 8

TiMo用动态注意力捕捉卫星影像时空变化,提升环境监测精度。

TiMo: Spatiotemporal Foundation Model for Satellite Image Time Series

  • 设计时空陀螺注意力,动态建模多尺度时空关系
  • 在百万级卫星影像数据上预训练,覆盖100,000个地点5年变化
  • 适用于森林砍伐、洪水检测等多类时空任务

卫星影像时序序列(SITS)提供了地球表面的持续观测,对环境管理与灾害评估至关重要。然而,现有时空基础模型依赖普通视觉变换器,未显式捕捉地物在时间和空间上的多尺度关系,限制了下游任务表现。为此,我们提出专用于SITS分析的层级视觉变换器模型TiMo。核心是引入时空陀螺注意力机制,动态捕捉时空演化中的多尺度模式。预训练阶段构建了百万级数据集MillionST,包含10万个地理区域、每个区域5年中10个时间阶段的影像,涵盖多样地理变化与季节性波动。利用该数据集,采用掩码图像建模进行预训练,使模型学习通用时空表示。在多个时空任务(包括森林砍伐监测、土地覆盖分割、作物类型分类、洪水检测)上的实验证明,TiMo优于现有先进方法。代码、模型与数据集将开源。

原文摘要 · Abstract (English)

Satellite image time series (SITS) provide continuous observations of the Earth's surface, making them essential for applications such as environmental management and disaster assessment. However, existing spatiotemporal foundation models rely on plain vision transformers, which encode entire temporal sequences without explicitly capturing multiscale spatiotemporal relationships between land objects. This limitation hinders their effectiveness in downstream tasks. To overcome this challenge, we propose TiMo, a novel hierarchical vision transformer foundation model tailored for SITS analysis. At its core, we introduce a spatiotemporal gyroscope attention mechanism that dynamically captures evolving multiscale patterns across both time and space. For pre-training, we curate MillionST, a large-scale dataset of one million images from 100,000 geographic locations, each captured across 10 temporal phases over five years, encompassing diverse geospatial changes and seasonal variations. Leveraging this dataset, we adapt masked image modeling to pre-train TiMo, enabling it to effectively learn and encode generalizable spatiotemporal representations.Extensive experiments across multiple spatiotemporal tasks-including deforestation monitoring, land cover segmentation, crop type classification, and flood detection-demonstrate TiMo's superiority over state-of-the-art methods. Code, model, and dataset will be released at https://github.com/MiliLab/TiMo.

卫星影像时空建模视觉变换器预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。