用视觉文本大模型提升不规则时间序列预测的上下文理解能力
MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs
- 通过多模态大模型生成图像和文本,捕捉复杂时间模式
- 在真实数据集上比现有方法最高提升12.3%的预测精度
- 适合需要理解时序上下文的工业与金融场景
不规则采样时间序列(ISTS)在现实场景中广泛存在,不同变量以异步、非均匀间隔观测。现有方法仅依赖历史观测进行预测,难以学习上下文语义和细粒度时间模式。为此,我们提出MM-ISTS,一个融合视觉-文本大语言模型的多模态预测框架,连接时间、视觉与文本模态。该框架采用两阶段编码:首先通过跨模态视觉-文本编码模块自动生成信息丰富的视觉图像与文本数据,配合多模态大模型(MLLMs)实现对复杂时间模式与完整上下文的理解;其次,基于不规则时间序列编码提取互补且丰富的时序特征,包含多视图嵌入融合与时间-变量编码器。此外,设计自适应查询特征提取器压缩MLLM token嵌入,降低计算开销;并引入模态感知门控的多模态对齐模块,缓解模态差异。在真实数据集上的大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across diverse variables. Existing ISTS forecasting methods often solely utilize historical observations to predict future ones while falling short in learning contextual semantics and fine-grained temporal patterns. To address these problems, we propose MM-ISTS, a multimodal ISTS forecasting framework augmented by vision-text large language models, which bridges temporal, visual, and textual modalities. MM-ISTS encompasses a two-stage encoding mechanism. In particular, a Cross-Modal Vision-Text Encoding module is proposed to automatically generate informative visual images and textual data, enabling the capture of intricate temporal patterns and comprehensive contextual understanding, in collaboration with multimodal LLMs (MLLMs). In parallel, ISTS encoding extracts complementary yet enriched temporal features from historical ISTS observations, including multi-view embedding fusion and a Temporal-Variable Encoder. Further, we propose an Adaptive Query-Based Feature Extractor to compress MLLM token embeddings while preserving useful knowledge, which in turn reduces computational costs. In addition, a Multimodal Alignment module with Modality-Aware Gating is designed to alleviate the modality gaps. Extensive experiments on real data offer insight into the effectiveness of the proposed solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。