针对多变量长时序预测,提出按深度分工的专家模型,性能更强且更省资源。
MoDEx: Mixture of Depth-specific Experts for Multivariate Long-term Time Series Forecasting
- 按层敏感性分析发现不同深度专注不同时间动态,据此设计分层专家结构。
- 在7个真实数据集上达顶尖水平,78%情况下排名第一,参数与计算量显著降低。
- 可无缝接入Transformer,提升其性能,适合追求高效高精度的时序任务应用。
多变量长时序预测(LTSF)支撑交通流管理、太阳能调度和电力变压器监控等关键应用。现有LTSF范式采用嵌入、主干精炼和长程预测三阶段流程,但各主干层的行为仍不明确。本文提出层敏感性,一种基于梯度的度量方法,受GradCAM和有效感受野理论启发,量化每个时间点对层特征的正负贡献。应用于三层MLP主干后,发现其在建模输入序列时间动态上存在深度特异性分工。受此启发,提出轻量级分层专家混合模型MoDEx,以深度特异性MLP专家替代复杂主干。MoDEx在七个真实世界基准上达到领先精度,在78%的场景中排名第一,同时显著减少参数量与计算开销。该模型还可无缝集成至Transformer变体中,持续提升性能,展现出作为高效高精度LTSF框架的强大泛化能力。
原文摘要 · Abstract (English)
Multivariate long-term time series forecasting (LTSF) supports critical applications such as traffic-flow management, solar-power scheduling, and electricity-transformer monitoring. The existing LTSF paradigms follow a three-stage pipeline of embedding, backbone refinement, and long-horizon prediction. However, the behaviors of individual backbone layers remain underexplored. We introduce layer sensitivity, a gradient-based metric inspired by GradCAM and effective receptive field theory, which quantifies both positive and negative contributions of each time point to a layer's latent features. Applying this metric to a three-layer MLP backbone reveals depth-specific specialization in modeling temporal dynamics in the input sequence. Motivated by these insights, we propose MoDEx, a lightweight Mixture of Depth-specific Experts, which replaces complex backbones with depth-specific MLP experts. MoDEx achieves state-of-the-art accuracy on seven real-world benchmarks, ranking first in 78 percent of cases, while using significantly fewer parameters and computational resources. It also integrates seamlessly into transformer variants, consistently boosting their performance and demonstrating robust generalizability as an efficient and high-performance LTSF framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。