arXiv:2506.14786cs.LGcs.AI2025-06NeurIPS被引 2

让卫星图像与时间序列对齐,提升台风预测精度12%。

PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series

  • 用物理规律设计位置编码,把时空信息嵌入视觉语言模型。
  • 在最大公开卫星数据集上,台风强度预测准确率提升12%。
  • 适合气候建模、遥感分析等需融合图像与时间数据的场景。

多模态时间序列预测在多个领域具有基础性作用,如利用卫星图像和数值数据预测台风。然而现有方法主要依赖文本辅助时间序列预测,忽视了时间序列数据中的视觉信息。同时,模型难以有效捕捉卫星图像中蕴含的物理信息,如时间和地理空间上下文。为此,我们提出物理感知位置编码(PIPE),一种轻量级方法,将物理信息嵌入视觉语言模型(VLMs)。PIPE引入两项关键创新:(1) 物理感知的位置索引机制,将物理规律映射为位置标识;(2) 变体频率位置编码,用于在嵌入空间中编码物理变量的频率特征及标记序列顺序。通过保留物理信息与时序结构,PIPE显著提升多模态对齐与预测精度。在最具代表性且规模最大的开源卫星图像数据集上,该方法在深度学习与气候领域方法中均达到领先水平,多项基准测试表现优异,台风强度预测较先前工作提升12%。代码见补充材料。

原文摘要 · Abstract (English)

Multimodal time series forecasting is foundational in various fields, such as utilizing satellite imagery and numerical data for predicting typhoons in climate science. However, existing multimodal approaches primarily focus on utilizing text data to help time series forecasting, leaving the visual data in existing time series datasets untouched. Furthermore, it is challenging for models to effectively capture the physical information embedded in visual data, such as satellite imagery's temporal and geospatial context, which extends beyond images themselves. To address this gap, we propose physics-informed positional encoding (PIPE), a lightweight method that embeds physical information into vision language models (VLMs). PIPE introduces two key innovations: (1) a physics-informed positional indexing scheme for mapping physics to positional IDs, and (2) a variant-frequency positional encoding mechanism for encoding frequency information of physical variables and sequential order of tokens within the embedding space. By preserving both the physical information and sequential order information, PIPE significantly improves multimodal alignment and forecasting accuracy. Through the experiments on the most representative and the largest open-sourced satellite image dataset, PIPE achieves state-of-the-art performance in both deep learning forecasting and climate domain methods, demonstrating superiority across benchmarks, including a 12% improvement in typhoon intensity forecasting over prior works. Our code is provided in the supplementary material.

多模态卫星图像时间序列物理编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。