Aurora让时间序列预测能跨领域通用,支持图文输入和零样本推理。
Aurora: Towards Universal Generative Multimodal Time Series Forecasting
- 用多模态提示引导时序建模,自动聚焦文本图像中的领域知识。
- 在5个基准上均达最好效果,多模态场景下性能领先。
- 适合需要跨领域预测、有图文数据的场景,如金融与医疗分析。
时间序列预测中跨领域泛化至关重要,因相同历史信息可能因领域特性导致不同未来趋势。现有方法或仅构建单模态基础模型,忽视文本等模态中的领域知识;或为端到端监督模型,不支持零样本跨领域推理。本文提出Aurora,一个支持多模态输入与零样本推理的时序基础模型。其在跨领域多模态时序语料库上预训练,通过分词、编码与蒸馏,自适应提取文本或图像中的关键领域知识,并利用模态引导的多头自注意力机制将其注入时序表征。解码阶段,多模态表示生成未来标记的条件与原型,结合新型原型引导流匹配实现生成式概率预测。在TimeMMD、TSFM-Bench、ProbTS、TFB和EPF共5个知名基准上的实验表明,Aurora在单模态与多模态场景下均保持一致领先性能。
原文摘要 · Abstract (English)
Cross-domain generalization is very important in Time Series Forecasting because similar historical information may lead to distinct future trends due to the domain-specific characteristics. Recent works focus on building unimodal time series foundation models and end-to-end multimodal supervised models. Since domain-specific knowledge is often contained in modalities like texts, the former lacks the explicit utilization of them, thus hindering the performance. The latter is tailored for end-to-end scenarios and does not support zero-shot inference for cross-domain scenarios. In this work, we introduce Aurora, a Multimodal Time Series Foundation Model, which supports multimodal inputs and zero-shot inference. Pretrained on Cross-domain Multimodal Time Series Corpus, Aurora can adaptively extract and focus on key domain knowledge contained in corresponding text or image modalities, thus possessing strong cross-domain generalization capability. Through tokenization, encoding, and distillation, Aurora can extract multimodal domain knowledge as guidance and then utilizes a Modality-Guided Multi-head Self-Attention to inject them into the modeling of temporal representations. In the decoding phase, the multimodal representations are used to generate the conditions and prototypes of future tokens, contributing to a novel Prototype-Guided Flow Matching for generative probabilistic forecasting. Comprehensive experiments on 5 well-recognized benchmarks, including TimeMMD, TSFM-Bench, ProbTS, TFB, and EPF, demonstrate the consistent state-of-the-art performance of Aurora on both unimodal and multimodal scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。