自适应专家数量的高精度时序预测模型
Ada-MoGE: Adaptive Mixture of Gaussian Expert Model for Time Series Forecasting
- 根据频谱特征动态调整专家数量,匹配数据变化
- 6个基准测试均达顶尖性能,参数仅20万
- 适合需要高精度时序建模的工业与金融场景
多变量时序预测广泛应用于工业、交通和金融等领域。然而,时序数据的主导频率会随频谱分布演变而变化。传统混合专家(MoE)模型采用固定专家数量,难以适应此类变化,导致频率覆盖不均:专家过少会遗漏关键信息,过多则引入噪声。为此,我们提出Ada-MoGE,一种基于频谱强度与频率响应自适应确定专家数量的高斯混合专家模型,确保与输入数据的频谱分布对齐。该方法既避免因专家不足造成的信息丢失,也防止因专家过多带来的噪声污染。此外,为防止直接截断频带引入噪声,采用高斯带通滤波器平滑分解频域特征,进一步优化特征表示。实验表明,该模型在6个公开基准上实现顶尖性能,仅需0.2百万参数。
原文摘要 · Abstract (English)
Multivariate time series forecasts are widely used, such as industrial, transportation and financial forecasts. However, the dominant frequencies in time series may shift with the evolving spectral distribution of the data. Traditional Mixture of Experts (MoE) models, which employ a fixed number of experts, struggle to adapt to these changes, resulting in frequency coverage imbalance issue. Specifically, too few experts can lead to the overlooking of critical information, while too many can introduce noise. To this end, we propose Ada-MoGE, an adaptive Gaussian Mixture of Experts model. Ada-MoGE integrates spectral intensity and frequency response to adaptively determine the number of experts, ensuring alignment with the input data's frequency distribution. This approach prevents both information loss due to an insufficient number of experts and noise contamination from an excess of experts. Additionally, to prevent noise introduction from direct band truncation, we employ Gaussian band-pass filtering to smoothly decompose the frequency domain features, further optimizing the feature representation. The experimental results show that our model achieves state-of-the-art performance on six public benchmarks with only 0.2 million parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。