arXiv:2602.05646cs.LG2026-02被引 1

用多模态预训练提升时间序列分析能力,效果超越现有模型。

Empowering Time Series Analysis with Large-Scale Multimodal Pretraining

  • 构建多模态时间序列预训练框架,融合图像、文本与新闻数据。
  • 在包含十亿点的大型数据集上训练,零样本预测与异常检测达领先水平。
  • 适合需要跨领域泛化的时间序列研究者和工业应用开发者。

现有时间序列基础模型主要依赖大规模单模态预训练,缺乏互补模态以增强理解。构建多模态基础模型是自然演进方向,但面临两大挑战:一是缺乏统一的多模态预训练范式及大规模时间序列多模态语料库;二是如何有效融合异构模态并提升模型泛化能力。为此,我们首次提出面向时间序列分析的多模态预训练范式,利用内生模态(衍生图像与文本)和外生知识(真实世界新闻),提供时间序列的多视角理解。为支持该范式,我们开发了自动化数据构建流程,创建了首个涵盖六个领域的大型多模态时间序列数据集MM-TS,最多包含一亿个时间点。随后提出HORAI,一种频域增强型多模态基础模型,包含频域增强跨模态编码器和时频解码器,可高效融合多模态特征并提升跨模态与跨领域的泛化能力。在MM-TS上预训练后,HORAI在时间序列预测与异常检测任务上实现最先进的零样本性能,验证了其强大的泛化能力。

原文摘要 · Abstract (English)

While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural next step, but it faces key challenges: 1) lack of a unified multimodal pretraining paradigm and large-scale multimodal corpora for time series analysis; 2) how to effectively integrate heterogeneous modalities and enhance model generalization. To address these challenges, we take an early step toward multimodal foundation models for time series analysis. We first propose a multimodal pretraining paradigm that leverages time series with endogenous modalities (derived images and text) and exogenous knowledge (real-world news), providing a comprehensive multi-view perspective for time series analysis. To support this, we develop an automated data construction pipeline to curate MM-TS, the first large-scale multimodal time series dataset spanning six domains, with up to one billion points. Then we propose HORAI, a frequency-enhanced multimodal foundation model. It integrates two core components: the Frequency-enhanced Cross-Modality Encoder and the Time-Frequency Decoder, designed to effectively fuse multimodal features and enhance model generalization across modalities and domains. After pretraining on MM-TS, HORAI achieves state-of-the-art zero-shot performance on time series forecasting and anomaly detection tasks, demonstrating strong generalization.

时间序列多模态预训练泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。