用图结构融合多模态时序特征,提升预测精度与适应性
MGTS-Net: Exploring Graph-Enhanced Multimodal Fusion for Augmented Time Series Forecasting
- 构建异构图建模模态内依赖与跨模态对齐关系
- 动态加权融合长短周期预测结果,适应多尺度特征
- 在多个数据集上超越现有模型,轻量高效
近期时序预测研究尝试融合多模态特征以提升精度,但受限于细粒度时序模式提取不足、多模态信息融合不优以及对动态多尺度特征适应性差三大挑战。为此,我们提出MGTS-Net:一种用于时序预测的多模态图增强网络。模型包含三个核心模块:(1) 多模态特征提取层(MFE),根据时序、视觉和文本模态特性优化特征编码器,提取细粒度时序特征;(2) 多模态融合层(MFF),构建异构图以建模模态内时序依赖与跨模态对齐关系,并动态聚合多模态知识;(3) 多尺度预测层(MSP),通过动态加权融合短、中、长期预测器输出,自适应多尺度特征。大量实验表明,MGTS-Net在保持轻量化与高效率的同时,显著优于现有先进基线模型,验证了方法的有效性。
原文摘要 · Abstract (English)
Recent research in time series forecasting has explored integrating multimodal features into models to improve accuracy. However, the accuracy of such methods is constrained by three key challenges: inadequate extraction of fine-grained temporal patterns, suboptimal integration of multimodal information, and limited adaptability to dynamic multi-scale features. To address these problems, we propose MGTS-Net, a Multimodal Graph-enhanced Network for Time Series forecasting. The model consists of three core components: (1) a Multimodal Feature Extraction layer (MFE), which optimizes feature encoders according to the characteristics of temporal, visual, and textual modalities to extract temporal features of fine-grained patterns; (2) a Multimodal Feature Fusion layer (MFF), which constructs a heterogeneous graph to model intra-modal temporal dependencies and cross-modal alignment relationships and dynamically aggregates multimodal knowledge; (3) a Multi-Scale Prediction layer (MSP), which adapts to multi-scale features by dynamically weighting and fusing the outputs of short-term, medium-term, and long-term predictors. Extensive experiments demonstrate that MGTS-Net exhibits excellent performance with light weight and high efficiency. Compared with other state-of-the-art baseline models, our method achieves superior performance, validating the superiority of the proposed methodology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。