arXiv:2506.07411cs.AI2025-06被引 2

用大模型+强化学习实现云AI系统故障自动修复,恢复速度提升37%。

An Intelligent Fault Self-Healing Mechanism for Cloud AI Systems via Integration of Large Language Models and Deep Reinforcement Learning

  • 大模型解析日志语义,强化学习优化修复策略,双阶段协同工作。
  • 在未知故障场景下,恢复时间缩短37%,显著优于传统方法。
  • 适合关注智能运维、云系统可靠性的研发与工程人员。

随着云上AI系统规模与复杂度持续增长,故障检测与自适应恢复已成为保障服务可靠性与连续性的核心挑战。本文提出一种智能故障自愈机制(IFSHM),融合大语言模型(LLM)与深度强化学习(DRL),构建具备语义理解与策略优化能力的故障恢复框架。基于传统DRL控制模型,该方法采用两阶段混合架构:(1) 基于LLM的故障语义解析模块,可动态从多源日志与系统指标中提取深层上下文语义,精准识别潜在故障模式;(2) 基于强化学习的恢复策略优化器,学习云环境中故障类型与响应行为间的动态匹配关系。其创新点在于引入LLM进行环境建模与动作空间抽象,大幅提升强化学习的探索效率与泛化能力。同时,引入记忆引导的元控制器,结合强化学习回放与LLM提示微调策略,实现对新故障模式的持续适应并避免灾难性遗忘。在云故障注入平台上的实验表明,相较于现有DRL与规则方法,该IFSHM框架在未知故障场景下将系统恢复时间缩短37%。

原文摘要 · Abstract (English)

As the scale and complexity of cloud-based AI systems continue to increase, the detection and adaptive recovery of system faults have become the core challenges to ensure service reliability and continuity. In this paper, we propose an Intelligent Fault Self-Healing Mechanism (IFSHM) that integrates Large Language Model (LLM) and Deep Reinforcement Learning (DRL), aiming to realize a fault recovery framework with semantic understanding and policy optimization capabilities in cloud AI systems. On the basis of the traditional DRL-based control model, the proposed method constructs a two-stage hybrid architecture: (1) an LLM-driven fault semantic interpretation module, which can dynamically extract deep contextual semantics from multi-source logs and system indicators to accurately identify potential fault modes; (2) DRL recovery strategy optimizer, based on reinforcement learning, learns the dynamic matching of fault types and response behaviors in the cloud environment. The innovation of this method lies in the introduction of LLM for environment modeling and action space abstraction, which greatly improves the exploration efficiency and generalization ability of reinforcement learning. At the same time, a memory-guided meta-controller is introduced, combined with reinforcement learning playback and LLM prompt fine-tuning strategy, to achieve continuous adaptation to new failure modes and avoid catastrophic forgetting. Experimental results on the cloud fault injection platform show that compared with the existing DRL and rule methods, the IFSHM framework shortens the system recovery time by 37% with unknown fault scenarios.

故障自愈大模型强化学习云系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。