arXiv:2606.03184q-fin.CPcs.LG2026-06KDD

构建金融时间序列合成基准,精准诊断模型在波动、极端事件等场景下的表现差异。

FinStressTS: A Parametric Synthetic Benchmark for Time-Series Forecasting in Finance

  • 设计30个可控环境,覆盖六类金融机制,实现故障归因可视化。
  • 发现线性模型在波动与跳变场景中常优于复杂神经网络,尤其数据稀缺时。
  • 揭示分布对齐重要性:概率模型在非平稳或多峰分布下需更灵活架构。

金融预测因信号噪声比低、潜变量、重尾分布、状态切换和跳跃特征而困难。真实数据仅呈现单一实现路径,难以评估尾部风险校准或数据效率。本文提出FinStressTS,一个机制感知的合成基准,将模型表现与可控制的结构成因关联。该基准包含围绕六类机制(波动聚集、多尺度持久性、重尾冲击、状态切换、自激跳跃、零膨胀过程)的30个诊断环境。评估两点任务:点预测(用五种设定下的NMAE)和概率预测(用已知生成机制下的CRPS)。测试15种模型,涵盖经典方法(HAR、VAR)、Transformer架构(PatchTST、iTransformer)及深度概率模型(DeepAR、TSFlow),并使用学习曲线衡量样本效率。结果揭示三方面洞见:一、性能高度依赖机制类型,自回归与线性模型在波动、重尾与跳跃驱动环境中表现优异,常超越Transformer;二、分布对齐至关重要,参数化概率模型在平稳环境下校准良好,而灵活模型在多模态或稀疏分布下更有优势;三、神经模型通常需要更多数据才能追平简单基线,在学习潜状态或复杂分布时增益显著。FinStressTS为诊断失败模式与推动风险敏感预测提供开放框架。

原文摘要 · Abstract (English)

Financial forecasting is difficult due to low signal-to-noise ratios, latent factors, heavy tails, regime shifts, and jumps. Real-world benchmarks offer limited failure attribution: researchers can observe underperformance, but often cannot isolate why because mechanisms are unobservable and entangled. Real financial data reveal only one realized path, making it difficult to assess tail-risk calibration or data efficiency. We introduce FinStressTS, a mechanism-aware synthetic benchmark that links model behavior to controlled structural causes. FinStressTS comprises 30 diagnostic environments around six mechanism families: volatility clustering, multi-scale persistence, heavy-tailed shocks, regime switching, self-exciting jumps, and zero-inflated processes. We evaluate two tasks: point forecasting, using NMAE across five settings, and probabilistic forecasting, using CRPS under known data-generating mechanisms. We benchmark 15 models, from classical methods (HAR, VAR) to Transformer forecasters (PatchTST, iTransformer) and deep probabilistic architectures (DeepAR, TSFlow), and use learning curves to measure sample efficiency. Our evaluation reveals three insights. First, performance is mechanism-dependent: autoregressive and linear models are highly competitive, and often outperform Transformer-based models, in several volatility-, tail-, and jump-driven environments. Second, distributional alignment matters: parametric probabilistic models such as DeepAR calibrate well in stationary settings, while flexible models can help when distributions become multimodal or sparse. Third, neural models often require more data to match simple baselines, with larger gains mainly when learning latent regimes or complex distributions. FinStressTS provides an open framework for diagnosing failure modes and advancing risk-aware forecasting.

时间序列金融建模合成数据风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。