用大模型自研智能诊断系统,提升AI集群故障处理效率
Enhancing Cluster Resilience: LLM-agent Based Autonomous Intelligent Cluster Diagnosis System and Evaluation Framework
- 基于大模型与检索增强生成技术构建自主诊断代理
- 在多维度测试中诊断准确率显著优于传统方法
- 适合需要高可靠性的AI集群运维团队使用
大型语言模型(LLM)及相关技术如检索增强生成(RAG)和思维图(DoT)的进步,使具备自主集群诊断与排错能力的智能系统成为可能。通过将这些技术与自对弈方法结合,我们开发了一种基于LLM代理的系统,可自主诊断并修复AI集群中的问题。创新包括专用于集群诊断的知识库、优化的LLM算法、实际部署策略以及针对该领域评估LLM能力的基准。在多个维度的大量实验中,证明了本系统在检测和修复性能问题方面比传统方法更高效、更准确。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) and related technologies such as Retrieval-Augmented Generation (RAG) and Diagram of Thought (DoT) have enabled the creation of autonomous intelligent systems capable of performing cluster diagnostics and troubleshooting. By integrating these technologies with self-play methodologies, we have developed an LLM-agent system designed to autonomously diagnose and resolve issues within AI clusters. Our innovations include a knowledge base tailored for cluster diagnostics, enhanced LLM algorithms, practical deployment strategies for agents, and a benchmark specifically designed for evaluating LLM capabilities in this domain. Through extensive experimentation across multiple dimensions, we have demonstrated the superiority of our system in addressing the challenges faced in cluster diagnostics, particularly in detecting and rectifying performance issues more efficiently and accurately than traditional methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。