综述大模型在运维智能中的应用现状与挑战
A Survey of AIOps in the Era of Large Language Models
- 分析183篇论文,梳理大模型处理故障数据的多种方法
- 发现新运维任务涌现,传统任务研究趋于饱和
- 揭示评估体系不统一,适合研究者和从业者参考
随着大语言模型(LLMs)日益复杂和普及,其在人工智能运维(AIOps)各类任务中的应用受到广泛关注。然而,对大模型在该领域的影响、潜力及局限性的系统性理解仍处于初期阶段。为填补这一空白,我们开展了一项关于LLM4AIOps的深度调研,聚焦大模型如何优化流程并提升运维效果。基于2020年1月至2024年12月期间发表的183篇文献,回答了四个核心研究问题(RQs)。RQ1分析了多样化的故障数据源,包括基于大模型的旧数据处理技术和由大模型催生的新数据来源;RQ2探讨了AIOps任务的演进,揭示了新任务的出现及各任务的发表趋势;RQ3研究了应用于解决运维挑战的大模型方法;RQ4回顾了针对集成大模型的AIOps方法的评估方法。基于上述发现,我们讨论了当前技术进展与趋势,识别现有研究的不足,并提出了未来探索的有前景方向。
原文摘要 · Abstract (English)
As large language models (LLMs) grow increasingly sophisticated and pervasive, their application to various Artificial Intelligence for IT Operations (AIOps) tasks has garnered significant attention. However, a comprehensive understanding of the impact, potential, and limitations of LLMs in AIOps remains in its infancy. To address this gap, we conducted a detailed survey of LLM4AIOps, focusing on how LLMs can optimize processes and improve outcomes in this domain. We analyzed 183 research papers published between January 2020 and December 2024 to answer four key research questions (RQs). In RQ1, we examine the diverse failure data sources utilized, including advanced LLM-based processing techniques for legacy data and the incorporation of new data sources enabled by LLMs. RQ2 explores the evolution of AIOps tasks, highlighting the emergence of novel tasks and the publication trends across these tasks. RQ3 investigates the various LLM-based methods applied to address AIOps challenges. Finally, RQ4 reviews evaluation methodologies tailored to assess LLM-integrated AIOps approaches. Based on our findings, we discuss the state-of-the-art advancements and trends, identify gaps in existing research, and propose promising directions for future exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。