根据数据变化时机重训,能省成本且保持预测精度。
Prepared for the Unknown: Adapting AIOps Capacity Forecasting Models to Data Changes
- 按数据漂移检测决定是否重训,而非固定周期
- 多数情况下游漂移重训精度接近定期重训
- 数据快速变化时仍推荐定期重训以保准度
容量管理对软件组织有效分配资源、满足运营需求至关重要。预测未来资源需求通常依赖数据驱动的分析和机器学习(ML)预测模型,这些模型需频繁重训以适应不断演变的数据。持续重训成本高且难扩展,给工程团队带来准确率与效率的平衡难题。仅在数据变化时重训看似更高效,但其对精度的影响尚待验证。本文研究基于数据变化检测的漂移驱动重训与定期重训在时间序列容量预测中的效果。结果表明,在大多数情况下,漂移驱动重训可达到与定期重训相当的预测精度,是一种更具成本效益的策略。但在数据变化迅速的情况下,定期重训仍更优,能最大化预测精度。研究为软件团队优化预测系统提供了可操作的洞见,可在降低重训开销的同时维持稳健性能。
原文摘要 · Abstract (English)
Capacity management is critical for software organizations to allocate resources effectively and meet operational demands. An important step in capacity management is predicting future resource needs often relies on data-driven analytics and machine learning (ML) forecasting models, which require frequent retraining to stay relevant as data evolves. Continuously retraining the forecasting models can be expensive and difficult to scale, posing a challenge for engineering teams tasked with balancing accuracy and efficiency. Retraining only when the data changes appears to be a more computationally efficient alternative, but its impact on accuracy requires further investigation. In this work, we investigate the effects of retraining capacity forecasting models for time series based on detected changes in the data compared to periodic retraining. Our results show that drift-based retraining achieves comparable forecasting accuracy to periodic retraining in most cases, making it a cost-effective strategy. However, in cases where data is changing rapidly, periodic retraining is still preferred to maximize the forecasting accuracy. These findings offer actionable insights for software teams to enhance forecasting systems, reducing retraining overhead while maintaining robust performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。