用大模型+残差结构分析微服务故障根源,效果更好更快。
Root Cause Analysis Method Based on Large Language Models with Residual Connection Structures
- 设计残差式层级融合结构,整合多源监控数据
- 在CCF-AIOps数据集上实现高准确率与高效分析
- 适合需要快速定位复杂系统故障的工程师
在复杂大规模微服务架构中,故障根源定位仍具挑战性。由于微服务间故障传播复杂、遥测数据(包括指标、日志、链路追踪)维度高,现有根因分析(RCA)方法效果受限。本文提出一种基于大语言模型(LLM)的残差连接式RCA方法(RC-LLM)。设计类残差的层次化融合结构,以集成多源遥测数据,并利用大语言模型的上下文推理能力,建模时间序列与跨微服务的因果依赖关系。在CCF-AIOps微服务数据集上的实验表明,RC-LLM在根因分析中表现出优异的准确性与效率。
原文摘要 · Abstract (English)
Root cause localization remain challenging in complex and large-scale microservice architectures. The complex fault propagation among microservices and the high dimensionality of telemetry data, including metrics, logs, and traces, limit the effectiveness of existing root cause analysis (RCA) methods. In this paper, a residual-connection-based RCA method using large language model (LLM), named RC-LLM, is proposed. A residual-like hierarchical fusion structure is designed to integrate multi-source telemetry data, while the contextual reasoning capability of large language models is leveraged to model temporal and cross-microservice causal dependencies. Experimental results on CCF-AIOps microservice datasets demonstrate that RC-LLM achieves strong accuracy and efficiency in root cause analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。