利用服务间因果关系提升网页流量预测准确率
Using Causality for Enhanced Prediction of Web Traffic Time Series
- 设计新模块CCMPlus捕捉多服务间的因果关系特征
- 在微软Azure等数据集上实现更低的MSE和MAE误差
- 适合需要精准流量预测的云服务运维人员
预测网络服务流量具有重要社会价值,可用于动态资源伸缩、负载均衡、异常检测、服务等级协议合规及欺诈检测等场景。网络服务流量随时间频繁剧烈波动,受异构用户行为影响,精准预测极具挑战。以往研究多采用统计方法和神经网络挖掘历史流量特征,但普遍忽视了服务间的因果关系。受生态系统中因果关系启发,我们实证识别出服务间的因果关联。为此提出神经网络模块CCMPlus,用于提取跨服务的因果特征,可无缝集成至现有时序模型中,持续提升预测性能。理论上证明CCMPlus生成的因果相关矩阵能有效捕捉服务间因果关系。在微软Azure、阿里巴巴集团及蚂蚁集团的真实数据集上验证,本方法在均方误差(MSE)和平均绝对误差(MAE)上均优于当前最优方法,证实了利用因果关系提升预测效果的有效性。
原文摘要 · Abstract (English)
Predicting web service traffic has significant social value, as it can be applied to various practical scenarios, including but not limited to dynamic resource scaling, load balancing, system anomaly detection, service-level agreement compliance, and fraud detection. Web service traffic is characterized by frequent and drastic fluctuations over time and are influenced by heterogeneous web user behaviors, making accurate prediction a challenging task. Previous research has extensively explored statistical approaches, and neural networks to mine features from preceding service traffic time series for prediction. However, these methods have largely overlooked the causal relationships between services. Drawing inspiration from causality in ecological systems, we empirically recognize the causal relationships between web services. To leverage these relationships for improved web service traffic prediction, we propose an effective neural network module, CCMPlus, designed to extract causal relationship features across services. This module can be seamlessly integrated with existing time series models to consistently enhance the performance of web service traffic predictions. We theoretically justify that the causal correlation matrix generated by the CCMPlus module captures causal relationships among services. Empirical results on real-world datasets from Microsoft Azure, Alibaba Group, and Ant Group confirm that our method surpasses state-of-the-art approaches in Mean Squared Error (MSE) and Mean Absolute Error (MAE) for predicting service traffic time series. These findings highlight the efficacy of leveraging causal relationships for improved predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。