用门控机制路由专家模型,提升股票波动率预测精度与训练稳定性
Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting

- 将状态变量仅用于专家路由,不直接参与预测
- 在1027只美股上,准确率和稳定性均优于对比模型
- 适合关注金融时变特征建模的量化研究者
金融波动率具有状态依赖性,但直接将状态信息引入神经网络可能引发训练不稳定。本文研究五日实现波动率对1027只美国股票的预测,在滚动前瞻评估框架下,保持模型容量、超参数调优和随机种子一致。提出RG-ResMoE:一种仅用状态变量进行专家路由的残差混合专家架构。基础模型从股票特征预测波动率,门控网络利用状态变量分配残差修正。该模型在主要美国样本中持续优于容量相当的MLP,预测准确性和训练稳定性均更优;日本独立面板也呈现类似优势。直接将状态变量接入输入端会降低性能与稳定性,而限制其仅用于门控则提升准确性与风险价值校准。硬路由始终劣于软路由。结果表明,在紧凑神经波动率模型中,混合专家的核心价值不在于扩大容量,而在于控制非平稳状态信息的影响路径。
原文摘要 · Abstract (English)
Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model. We study five-day realized-volatility forecasts for 1,027 U.S. equities using a rolling walk-forward evaluation framework in which information, model capacity, hyperparameter tuning, and random seeds are matched across architectures. We propose RG-ResMoE, a regime-gated residual mixture-of-experts architecture in which regime information is used only for expert routing rather than for direct forecasting. The base predictor models volatility from stock features, while a gating network uses regime state variables to route residual corrections. RG-ResMoE consistently outperforms a capacity-matched MLP in both forecasting accuracy and training stability in the main U.S. study. Similar gains are observed on an independent Japanese panel. The integration pathway is decisive: appending the same regime variables directly to the forecasting input degrades both predictive performance and training stability, whereas restricting them to the routing gate improves accuracy and Value-at-Risk calibration. Hard routing consistently underperforms soft routing. The results suggest that, in compact neural volatility forecasting models, the primary value of mixture-of-experts models lies less in increasing model capacity than in controlling how nonstationary regime information influences prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。