通过可控合成数据评估多变量长时序预测模型在不同噪声与频率下的表现。
Benchmarking M-LTSF: Frequency and Noise-Based Evaluation of Multivariate Long Time Series Forecasting Models
- 构建可调节的合成数据集,模拟真实时间序列中的信号、噪声与频率特性。
- 发现所有模型在无法捕捉完整周期时性能大幅下降,且对不同信号类型有偏好差异。
- 揭示S-Mamba和iTransformer在频谱重建上更优,适合特定噪声与信号场景。
多变量长时序预测(M-LTSF)模型的鲁棒性评估仍具挑战,因多数研究依赖噪声特性未知的真实数据集。本文提出一种基于仿真的评估框架,生成可参数化的合成数据集,每个实例对应不同的信号成分、噪声类型、信噪比及频率特征。该框架可系统化地在受控且多样化的场景下评估M-LTSF模型。我们对四种代表性架构——S-Mamba(状态空间)、iTransformer(基于Transformer)、R-Linear(线性)和Autoformer(分解式)进行基准测试。结果表明,当回溯窗口无法覆盖完整季节周期时,所有模型性能均显著下降;其中S-Mamba和Autoformer在锯齿波模式下表现最佳,而R-Linear和iTransformer更适应正弦信号。白噪声与布朗运动噪声在低信噪比下普遍降低模型性能,且S-Mamba对趋势-噪声组合敏感,iTransformer则对季节-噪声组合敏感。进一步频谱分析显示,S-Mamba和iTransformer具有更优的频率重建能力。本研究基于原理驱动的合成测试平台,通过均方误差(MSE)聚合,为模型选型提供了基于信号与噪声特性的具体指导。
原文摘要 · Abstract (English)
Understanding the robustness of deep learning models for multivariate long-term time series forecasting (M-LTSF) remains challenging, as evaluations typically rely on real-world datasets with unknown noise properties. We propose a simulation-based evaluation framework that generates parameterizable synthetic datasets, where each dataset instance corresponds to a different configuration of signal components, noise types, signal-to-noise ratios, and frequency characteristics. These configurable components aim to model real-world multivariate time series data without the ambiguity of unknown noise. This framework enables fine-grained, systematic evaluation of M-LTSF models under controlled and diverse scenarios. We benchmark four representative architectures S-Mamba (state-space), iTransformer (transformer-based), R-Linear (linear), and Autoformer (decomposition-based). Our analysis reveals that all models degrade severely when lookback windows cannot capture complete periods of seasonal patters in the data. S-Mamba and Autoformer perform best on sawtooth patterns, while R-Linear and iTransformer favor sinusoidal signals. White and Brownian noise universally degrade performance with lower signal-to-noise ratio while S-Mamba shows specific trend-noise and iTransformer shows seasonal-noise vulnerability. Further spectral analysis shows that S-Mamba and iTransformer achieve superior frequency reconstruction. This controlled approach, based on our synthetic and principle-driven testbed, offers deeper insights into model-specific strengths and limitations through the aggregation of MSE scores and provides concrete guidance for model selection based on signal characteristics and noise conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。