用大模型自动生成可解释的异常检测规则,提升云监控准确率
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
- 通过大模型自动生成可解释、可复现的异常规则作为中间表示
- 在公开和内部数据集上分别提升9.5%和28.3%的F1分数
- 适合需要高可信度异常检测的云平台运维团队
云基础设施可观测性对服务提供商至关重要,推动了异常检测系统的广泛应用。然而,现有系统难以同时实现可解释性、可复现性和自主性,而这三项是生产环境不可或缺的特性。我们提出Argos,一个基于大语言模型(LLM)的时序异常检测代理系统。Argos采用可解释且可复现的异常规则作为中间表示,并利用LLM实现规则的自主生成。系统通过多个协作智能体高效训练出无错误且精度有保障的规则,并部署用于低成本在线异常检测。实验表明,Argos在公开异常检测数据集和微软内部数据集上的F1分数分别提升最高达9.5%和28.3%,优于现有最先进方法。
原文摘要 · Abstract (English)
Observability in cloud infrastructure is critical for service providers, driving the widespread adoption of anomaly detection systems for monitoring metrics. However, existing systems often struggle to simultaneously achieve explainability, reproducibility, and autonomy, which are three indispensable properties for production use. We introduce Argos, an agentic system for detecting time-series anomalies in cloud infrastructure by leveraging large language models (LLMs). Argos proposes to use explainable and reproducible anomaly rules as intermediate representation and employs LLMs to autonomously generate such rules. The system will efficiently train error-free and accuracy-guaranteed anomaly rules through multiple collaborative agents and deploy the trained rules for low-cost online anomaly detection. Through evaluation results, we demonstrate that Argos outperforms state-of-the-art methods, increasing $F_1$ scores by up to $9.5\%$ and $28.3\%$ on public anomaly detection datasets and an internal dataset collected from Microsoft, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。