arXiv:2510.23118cs.CV2025-10

用统一表示空间和时间数据,让卫星图像生成全球温度变化

Quantizing Space and Time: Fusing Time Series and Images for Earth Observation

  • 将时序数据离散化后与图像对齐,构建跨模态统一表征
  • 在地球观测任务中,生成性能比专用融合方法高6%(R²)
  • 适合需要多模态融合的遥感、气候建模等研究者使用

我们提出一种面向时序与单时刻图像的无任务依赖多模态融合框架,支持跨模态生成并提升下游任务表现。该方法探索了时序数据的确定性与学习型量化策略,并引入掩码相关性学习目标,将离散的图像与时序标记对齐于统一表示空间。在地球观测领域实例化后,预训练模型能从卫星影像生成一致的全球温度分布,并通过反事实实验验证。在多个下游任务中,该无任务依赖预训练相比任务特定融合方法平均提升6%(R²)和2%(RMSE),优于基线方法50%(R²)和12%(RMSE)。最后,我们分析了不同模态的梯度敏感性,揭示模型鲁棒性。代码、数据与权重将开源。

原文摘要 · Abstract (English)

We propose a task-agnostic framework for multimodal fusion of time series and single timestamp images, enabling cross-modal generation and robust downstream performance. Our approach explores deterministic and learned strategies for time series quantization and then leverages a masked correlation learning objective, aligning discrete image and time series tokens in a unified representation space. Instantiated in the Earth observation domain, the pretrained model generates consistent global temperature profiles from satellite imagery and is validated through counterfactual experiments. Across downstream tasks, our task-agnostic pretraining outperforms task-specific fusion by 6% in R^2 and 2% in RMSE on average, and exceeds baseline methods by 50% in R^2 and 12% in RMSE. Finally, we analyze gradient sensitivity across modalities, providing insights into model robustness. Code, data, and weights will be released under a permissive license.

多模态融合遥感时序生成图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。