用代理模型+可预测性分析,让复杂预测模型既透明又可信。
Faithful and Interpretable Explanations for Complex Ensemble Time Series Forecasts using Surrogate Models and Forecastability Analysis
- 用LightGBM模仿AutoGluon预测,生成稳定可解释的特征贡献值。
- 谱可预测性越高,预测越准,解释也越可信,相关性达0.87。
- 适合需要信任预测结果的金融、供应链等场景用户使用。
现代时间序列预测越来越多依赖AutoGluon等AutoML系统生成的复杂集成模型,虽精度高,但透明度与可解释性差。本文提出双轨框架:首先构建基于代理模型的解释方法,训练LightGBM忠实模仿AutoGluon的预测结果,支持稳定的SHAP特征归因;通过特征注入实验验证,提取的SHAP值与真实影响高度一致。其次引入谱可预测性分析,通过对比时间序列与其纯噪声基准的频谱特性,量化其内在可预测性。在M5数据集上的实证表明,谱可预测性越高,不仅预测精度越高(相关性0.87),且代理模型与原模型的解释一致性也更强。该指标可作为置信度评分和过滤机制,帮助用户判断预测与解释的可靠性。研究还发现,按项目归一化对跨尺度序列生成有意义的SHAP解释至关重要。该框架实现了对前沿集成模型的实例级可解释性,并提供可靠的预测可信度指标。
原文摘要 · Abstract (English)
Modern time series forecasting increasingly relies on complex ensemble models generated by AutoML systems like AutoGluon, delivering superior accuracy but with significant costs to transparency and interpretability. This paper introduces a comprehensive, dual-approach framework that addresses both the explainability and forecastability challenges in complex time series ensembles. First, we develop a surrogate-based explanation methodology that bridges the accuracy-interpretability gap by training a LightGBM model to faithfully mimic AutoGluon's time series forecasts, enabling stable SHAP-based feature attributions. We rigorously validated this approach through feature injection experiments, demonstrating remarkably high faithfulness between extracted SHAP values and known ground truth effects. Second, we integrated spectral predictability analysis to quantify each series' inherent forecastability. By comparing each time series' spectral predictability to its pure noise benchmarks, we established an objective mechanism to gauge confidence in forecasts and their explanations. Our empirical evaluation on the M5 dataset found that higher spectral predictability strongly correlates not only with improved forecast accuracy but also with higher fidelity between the surrogate and the original forecasting model. These forecastability metrics serve as effective filtering mechanisms and confidence scores, enabling users to calibrate their trust in both the forecasts and their explanations. We further demonstrated that per-item normalization is essential for generating meaningful SHAP explanations across heterogeneous time series with varying scales. The resulting framework delivers interpretable, instance-level explanations for state-of-the-art ensemble forecasts, while equipping users with forecastability metrics that serve as reliability indicators for both predictions and their explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。