AdaMixT通过自适应融合多尺度专家模型,提升时间序列预测精度。
AdaMixT: Adaptive Weighted Mixture of Multi-Scale Expert Transformers for Time Series Forecasting
- 采用多尺度补丁与门控网络动态分配权重,实现灵活特征融合。
- 在8个基准数据集上均优于现有方法,尤其在复杂场景下表现突出。
- 适合需要高精度时序预测的工业、气象等实际应用。
多变量时间序列预测需基于历史观测值预测未来值。然而,现有方法多依赖预定义单尺度补丁或缺乏有效的多尺度特征融合机制,难以充分捕捉时间序列中的复杂模式,导致性能受限且泛化能力不足。为此,我们提出一种新型架构——自适应加权多尺度专家变换器(AdaMixT)。具体而言,AdaMixT引入多种尺度补丁,并结合通用预训练模型(GPM)与领域特定模型(DSM)进行多尺度特征提取。为应对时序特征的异质性,AdaMixT设计门控网络,动态分配不同专家的权重,实现自适应多尺度融合,从而提升预测准确性。在包含气象、交通、电力、流感及四个ETT数据集在内的8个常用基准上进行的全面实验,一致验证了AdaMixT在真实场景下的有效性。
原文摘要 · Abstract (English)
Multivariate time series forecasting involves predicting future values based on historical observations. However, existing approaches primarily rely on predefined single-scale patches or lack effective mechanisms for multi-scale feature fusion. These limitations hinder them from fully capturing the complex patterns inherent in time series, leading to constrained performance and insufficient generalizability. To address these challenges, we propose a novel architecture named Adaptive Weighted Mixture of Multi-Scale Expert Transformers (AdaMixT). Specifically, AdaMixT introduces various patches and leverages both General Pre-trained Models (GPM) and Domain-specific Models (DSM) for multi-scale feature extraction. To accommodate the heterogeneity of temporal features, AdaMixT incorporates a gating network that dynamically allocates weights among different experts, enabling more accurate predictions through adaptive multi-scale fusion. Comprehensive experiments on eight widely used benchmarks, including Weather, Traffic, Electricity, ILI, and four ETT datasets, consistently demonstrate the effectiveness of AdaMixT in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。