arXiv:2411.09047cs.LGcs.DC2024-11被引 26

IBM云实测数据集助力大规模系统异常检测研究

Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset

  • 基于4.5个月真实云监控数据,构建3.9万行、11.7万列高维数据集
  • 验证机器学习模型在复杂云系统中的异常检测效果
  • 适合云监控、系统可靠性与工业级数据研究者使用

随着大规模云系统日益复杂,有效的异常检测对保障系统可靠性和性能至关重要。然而,现有公开的、大规模真实数据集仍显不足。为此,本文介绍来自IBM云平台的真实数据集,覆盖4.5个月的IBM云控制台采集数据,包含39,365行、117,448列的遥测数据。同时,本文展示了机器学习模型在异常检测中的应用,并讨论了该过程中的关键挑战。本研究及配套数据集为云系统监控领域的研究人员和实践者提供了宝贵资源,有助于在真实数据上更高效地测试异常检测方法,推动鲁棒性解决方案的发展,以维护大规模云基础设施的健康与性能。

原文摘要 · Abstract (English)

As Large-Scale Cloud Systems (LCS) become increasingly complex, effective anomaly detection is critical for ensuring system reliability and performance. However, there is a shortage of large-scale, real-world datasets available for benchmarking anomaly detection methods. To address this gap, we introduce a new high-dimensional dataset from IBM Cloud, collected over 4.5 months from the IBM Cloud Console. This dataset comprises 39,365 rows and 117,448 columns of telemetry data. Additionally, we demonstrate the application of machine learning models for anomaly detection and discuss the key challenges faced in this process. This study and the accompanying dataset provide a resource for researchers and practitioners in cloud system monitoring. It facilitates more efficient testing of anomaly detection methods in real-world data, helping to advance the development of robust solutions to maintain the health and performance of large-scale cloud infrastructures.

异常检测云系统真实数据遥测数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。