用后验学习债务动态决定模型重训时机,更智能省成本。
Cost-sensitive retraining via posterior learning debt
- 以后验学习债务衡量模型过时程度,指导重训决策。
- 在72种场景中,债务阈值策略优于调优的固定周期重训。
- 适合长期部署的贝叶斯预测系统,尤其关注成本效率的人看。
已部署的预测系统通常按固定日历重训,但模型过时和重训负担随时间变化。本文将贝叶斯预测系统的重训建模为一个代价敏感的预测后悔决策问题。核心监控状态是后验学习债务,定义为参考影子后验与已部署冻结后验之间的Kullback--Leibler散度。决策层比较重训成本与等待的预期单期预测后悔。连续严重性版本在校准后的预期后悔超过重训成本时重训,而熟悉的两状态超额损失规则是其特例。实证研究在合成共轭模拟中进行了精确状态验证,包含冷启动的部署与影子正态逆高斯后验、独立更新/监控/评估批次、滞后部署动作、扩展基线网格及得分单位敏感性分析。在主要75%分位数得分单位缩放下,年龄调整的债务阈值策略在全部72个非稳定场景单元中优于调优的固定周期重训,在58个场景中优于调优的CUSUM,平均相对目标分别为0.677和0.975。债务效用与混合效用策略也显著优于调优的固定周期重训,但未全面超越调优的CUSUM。中位数与均值得分单位敏感性结果一致,而CUSUM比较仍依赖具体策略。贡献在于为已部署贝叶斯预测系统提供透明的决策层,而非漂移检测的通用替代方案。
原文摘要 · Abstract (English)
Deployed prediction systems are often retrained on fixed calendars, even when model staleness and retraining burden vary over time. This short communication formulates retraining for Bayesian prediction systems as a cost-sensitive predictive-regret decision. The central monitoring state is posterior learning debt, defined as the Kullback--Leibler divergence from a reference shadow posterior to the deployed frozen posterior. In the decision layer, a retraining cost is compared with the expected one-period predictive regret of waiting. A continuous-severity version retrains when calibrated expected regret exceeds the retraining cost, while the familiar two-state excess-loss rule is a special case. The empirical study is an exact-state proof-of-concept in a synthetic conjugate simulation with warm-started deployed and shadow normal-inverse-gamma posteriors, separate update, monitoring, and evaluation batches, lagged deployment actions, expanded baseline grids, and score-unit sensitivity. Under the primary 75th-percentile score-unit scaling, an age-adjusted debt-threshold policy improves on tuned calendar retraining in all 72 non-stable scenario cells and on tuned CUSUM in 58 of 72 cells, with mean relative objectives 0.677 and 0.975, respectively. Debt-utility and hybrid-utility policies also improve strongly over tuned calendar retraining, but they do not dominate tuned CUSUM. Median and mean score-unit sensitivities show the same main calendar result, while the CUSUM comparison remains policy-dependent. The contribution is a transparent decision layer for deployed Bayesian prediction systems, not a universal replacement for drift detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。