通过隐含市场状态审计,发现高准确率波动率模型在特定行情下仍会失效。
Latent-Regime Bias Auditing for Volatility Forecasting

- 基于训练数据聚类市场状态,构建可解释的隐含市场制度
- 发现部分模型在特定市场状态下存在严重尾部低估,误差达30%以上
- 适合风控人员和量化模型开发者评估模型可靠性边界
波动率预测通常用均方根误差(RMSE)和平均绝对误差(MAE)等整体指标评估,但这些指标可能掩盖对风险管理至关重要的条件性失败。本文提出一种模型无关的审计框架,用于检验波动率预测在隐含市场制度下的可靠性。我们利用仅基于训练数据的时间序列表示,将市场状态窗口聚类为若干制度,并在样本外分配制度标签,对比整体预测表现与制度条件下的偏差、尾部低估及敏感经济损失。应用于加密货币和ETF资产的日度波动率预测,结果显示:尽管某些模型具备竞争性整体精度,却在特定制度中表现出显著的制度性偏差和严重的尾部低估。研究建议,波动率预测应不仅看平均误差,更需关注预测何时何地失效。本框架将评估重点从‘哪个模型平均最准’转向‘哪些制度下看似准确的预测会失效’。
原文摘要 · Abstract (English)
Volatility forecasts are commonly evaluated with aggregate accuracy metrics such as RMSE and MAE, but these metrics can hide conditional failures that matter for risk management. This paper proposes a model-agnostic audit framework for evaluating whether volatility forecasts remain reliable across latent market regimes. We learn time-series representations of market-state windows, cluster them into regimes using only training information, assign regimes out of sample, and compare aggregate forecast behavior with regime-conditional bias, tail-underprediction, and underprediction-sensitive economic losses. Applied to daily volatility forecasting across cryptocurrency and ETF assets, the audit shows that models with competitive aggregate accuracy can still exhibit substantial regime-specific bias and severe tail underprediction. The results suggest that volatility forecasting should be evaluated not only by average error, but also by where and how forecasts become unreliable. Our framework shifts forecast evaluation from asking which model is most accurate on average to identifying the market regimes in which apparently accurate forecasts fail conditionally. Reproducibility: https://github.com/arthurchagas1/Latent-Regime-Bias-Auditing-for-Volatility-Forecasting
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。