arXiv:2510.03244cs.LGcs.AI2025-10

用视觉模型捕捉时间序列跨变量模式,提升预测精度

VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

  • 将多变量时间序列转为图像,利用预训练视觉模型感知空间关系
  • 双分支融合视觉与时间特征,跨模态注意力提升预测效果
  • 仅训练3.5%参数,适配资源有限但需高精度预测的场景

大型时间序列基础模型通常采用通道独立架构处理不同维度数据,但忽略了关键的跨通道依赖。现有跨模态方法主要依赖文本模态,未充分挖掘视觉模型在空间模式识别方面的潜力。为此,我们提出VFEM,一种利用预训练大视觉模型(LVMs)捕捉复杂跨变量模式的跨模态预测模型。VFEM将多变量时间序列转换为视觉表示,使LVM能够感知通道间未显式建模的空间关系。通过双分支结构,视觉与时间特征分别提取后经由跨模态注意力融合,实现两模态互补信息增强预测。仅冻结LVM并训练总参数的7.45%,在多个基准上取得有竞争力的表现,为多变量时间序列预测提供了新视角。

原文摘要 · Abstract (English)

Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores crucial cross-channel dependencies. Meanwhile, existing cross-modal methods predominantly rely on textual modalities, leaving the spatial pattern recognition capabilities of vision models underexplored for time series analysis. To address these limitations, we propose VFEM, a cross-modal forecasting model that leverages pre-trained large vision models (LVMs) to capture complex cross-variable patterns. VFEM transforms multivariate time series into visual representations, enabling LVMs to perceive spatial relationships that are not explicitly modeled by channel-independent models. Through a dual-branch architecture, visual and temporal features are independently extracted and then fused via cross-modal attention, allowing complementary information from both modalities to enhance forecasting. By freezing the LVM and training only 7.45% of the total parameters, VFEM achieves competitive performance on multiple benchmarks, offering a new perspective on multivariate time series forecasting.

时间序列预测跨模态融合视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。