arXiv:2604.02351cs.LG2026-04

提出动态可靠性控制框架,让模型部署更稳定且省成本。

Modeling and Controlling Deployment Reliability under Temporal Distribution Shift

论文配图:Modeling and Controlling Deployment Reliability under Temporal Distribution Shift
图 1 · 摘自论文原文
  • 把可靠性看作可追踪的动态状态,分判别与校准两方面。
  • 实验显示按漂移触发干预比持续重训更稳,成本降七成以上。
  • 适合金融等高风险场景,关注长期稳定性的实践者必看。

在非平稳环境中部署的机器学习模型会遭遇时间分布漂移,导致预测可靠性随时间下降。现有缓解策略如周期性重训练和校准通常只关注孤立时间点的平均性能,未显式建模可靠性在部署过程中的演化。本文提出一种以部署为中心的框架,将可靠性视为由判别力与校准度构成的动态状态。该状态在连续评估窗口中的轨迹可量化为波动性,使部署适应转化为多目标控制问题,需权衡可靠性稳定性与累积干预成本。在此框架下,定义了一类依赖状态的干预策略,并实证刻画了成本-波动性帕累托前沿。基于大规模时序信贷风险数据集(135万笔贷款,2007–2018年)的实验表明,选择性、漂移触发的干预策略相比持续滚动重训练,能实现更平滑的可靠性轨迹,同时显著降低运营成本。研究将时间漂移下的部署可靠性定位为可控的多目标系统,凸显策略设计在高风险表格应用中稳定性与成本权衡的关键作用。

原文摘要 · Abstract (English)

Machine learning models deployed in non-stationary environments are exposed to temporal distribution shift, which can erode predictive reliability over time. While common mitigation strategies such as periodic retraining and recalibration aim to preserve performance, they typically focus on average metrics evaluated at isolated time points and do not explicitly model how reliability evolves during deployment. We propose a deployment-centric framework that treats reliability as a dynamic state composed of discrimination and calibration. The trajectory of this state across sequential evaluation windows induces a measurable notion of volatility, allowing deployment adaptation to be formulated as a multi-objective control problem that balances reliability stability against cumulative intervention cost. Within this framework, we define a family of state-dependent intervention policies and empirically characterize the resulting cost-volatility Pareto frontier. Experiments on a large-scale, temporally indexed credit-risk dataset (1.35M loans, 2007-2018) show that selective, drift-triggered interventions can achieve smoother reliability trajectories than continuous rolling retraining while substantially reducing operational cost. These findings position deployment reliability under temporal shift as a controllable multi-objective system and highlight the role of policy design in shaping stability-cost trade-offs in high-stakes tabular applications.

可靠性部署优化时间漂移控制框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。