提出E3Former模型,实现云端负载精准在线预测,显著提升自动扩容效率。
Online Ensemble Transformer for Accurate Cloud Workload Forecasting in Predictive Auto-Scaling
- 构建多子网络协同的在线集成模型,动态适应高频率负载变化。
- 相比现有方法,平均预测误差降低10%,真实系统中资源利用率下降超40%。
- 已部署于字节跳动平台,支撑30+应用、超60万核资源自动伸缩。
在快速发展的云计算领域,无服务器系统对预测性自动扩容提出了更高要求,以应对固有的环境波动。其核心在于工作负载预测模型。现有模型难以快速适应在线负载流的变化,且难捕捉细粒度高频任务带来的复杂周期性。为此,我们提出新型在线集成模型E3Former,通过多个子网络协同,突破单一模型局限,显著提升预测精度与鲁棒性,同时计算开销极低,符合无服务器系统的轻量化理念。在真实负载数据集上的实验表明,该方法在在线预测任务中平均误差降低10%;在实际系统中的预测自动扩容测试进一步验证了其有效性。目前,该方法已部署于字节跳动智能水平Pod自动扩容(IHPA)平台,支撑包括抖音电商、头条、火山引擎在内的30余项应用,预测扩容能力覆盖超60万CPU核心。在基本保障服务质量的前提下,系统可降低资源利用率超过40%。
原文摘要 · Abstract (English)
In the swiftly evolving domain of cloud computing, the advent of serverless systems underscores the crucial need for predictive auto-scaling systems. This necessity arises to ensure optimal resource allocation and maintain operational efficiency in inherently volatile environments. At the core of a predictive auto-scaling system is the workload forecasting model. Existing forecasting models struggle to quickly adapt to the dynamics in online workload streams and have difficulty capturing the complex periodicity brought by fine-grained, high-frequency forecasting tasks. Addressing this, we propose a novel online ensemble model, E3Former, for online workload forecasting in large-scale predictive auto-scaling. Our model synergizes the predictive capabilities of multiple subnetworks to surmount the limitations of single-model approaches, thus ensuring superior accuracy and robustness. Remarkably, it accomplishes this with a minimal increase in computational overhead, adhering to the lean operational ethos of serverless systems. Through extensive experimentation on real-world workload datasets, we establish the efficacy of our ensemble model. In online forecasting tasks, the proposed method reduces forecast error by an average of 10%, and its effectiveness is further demonstrated through a predictive auto-scaling test in the real-life online system. Currently, our method has been deployed within ByteDance's Intelligent Horizontal Pod Auto-scaling (IHPA) platform, which supports the stable operation of over 30 applications, such as Douyin E-Comerce, TouTiao, and Volcano Engine. The predictive auto-scaling capacity reaching over 600,000 CPU cores. On the basis of essentially ensuring service quality, the predictive auto-scaling system can reduce resource utilization by over 40%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。