基于操作上下文的云系统异常检测框架,提升大规模监控准确性与早期发现能力。
ClouDens: Operational Context-Aware Anomaly Detection for Large-scale Cloud System Monitoring

- 按领域划分高维日志,构建上下文感知图模型捕捉服务依赖关系
- 在IBM云数据集上比GRU模型NAB得分更高,实现更早、更广覆盖的异常检测
- 适用于云运维团队及系统监控研发人员,尤其关注复杂依赖与稀疏数据场景
随着云计算基础设施规模与复杂性的快速增长,大规模云系统(LCS)的网络监控日益困难,亟需自动化且可靠的异常检测以保障服务可用性。现代LCS持续生成来自分布式云服务的遥测日志,形成高维多变量时间序列,反映系统运行状态。由于维度极高、分布式组件间依赖复杂,以及服务间歇性活跃导致的严重稀疏性,异常检测面临挑战。本文首先对IBM Cloud Console平台的遥测日志进行实证研究,提出针对LCS监控优化的异常检测框架ClouDens,利用日志模式中编码的操作上下文属性提升检测精度与异常早期识别能力。ClouDens将高维遥测日志划分为领域引导子集,构建上下文感知图以建模服务依赖关系,并采用时空图神经网络实现基于预测的异常检测。我们在最新发布的IBM Cloud Telemetry Dataset上评估了ClouDens,结果表明其在计数型遥测特征上取得更高的NAB分数,相比基于GRU的模型具备更优的准确率、更早的异常发现和更广的覆盖范围。研究进一步揭示,遥测特征子集划分、操作上下文建模、评分策略及稀疏性填补均显著影响检测性能,为设计与公平评估LCS监控异常检测方法提供实践指导。
原文摘要 · Abstract (English)
With the rapid growth of cloud computing infrastructures in scale and complexity, network monitoring for Large-scale Cloud Systems (LCSs) has become increasingly challenging, requiring automated and reliable anomaly detection to maintain service availability. Modern LCSs continuously generate telemetry logs from distributed cloud services, producing high-dimensional multivariate time series that capture system operations. Detecting anomalies in this context is difficult due to extreme dimensionality, complex dependencies among distributed components, and severe sparsity from intermittently active services. Taking these challenges into account, we first conduct an empirical study on telemetry logs from the IBM Cloud Console platform, and then propose ClouDens, an anomaly detection framework tailored to LCS monitoring that leverages operational-context attributes encoded in the telemetry log schema to improve detection accuracy and early identification of anomalies. ClouDens partitions high-dimensional telemetry logs into domain-guided subsets, constructs a context-aware graph modeling operational service dependencies, and employs Spatio-Temporal Graph Neural Networks for forecasting-based anomaly detection. We evaluate ClouDens on the recently released IBM Cloud Telemetry Dataset and provide practical insights into designing reliable anomaly detection solutions for LCS monitoring. Results show ClouDens achieves higher NAB scores in count-based telemetry features, indicating more accurate, earlier anomaly detection with broader coverage than a GRU-based model. Our study further reveals that telemetry feature subsets, operational-context modeling, scoring strategies, and sparsity imputation all substantially influence detection performance, offering practical guidance for designing and fairly benchmarking anomaly detection approaches for LCS monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。