arXiv:2603.21522cs.SEcs.AI2026-03中稿 · FSE'26-IVR被引 6

用历史故障模式提升多智能体系统故障管理效率

Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation

  • 通过推理轨迹表示学习,统一建模单智能体与跨智能体行为
  • 实测可在多智能体系统中实现秒级故障检测与自愈
  • 适合关注AI系统可靠性的研发与运维人员

基于大语言模型的多智能体系统(MAS)在软件设计中展现出强大的推理与协作能力。随着系统复杂度提升,高效故障管理对保障可靠性至关重要。现有方法依赖逐条推理分析,效率低下,且忽略历史故障模式,影响诊断精度。本文首次开展实证研究,验证利用历史故障模式提升故障管理的必要性与潜力。在此基础上,提出EAGER框架,基于推理轨迹表示,采用无监督推理范围对比学习,编码单智能体内部推理与跨智能体协同行为,实现基于历史故障知识的实时步骤级故障检测、诊断与反射式修复。在三个开源MAS上的初步评估表明EAGER有效,为可靠多智能体系统运行提供了新方向。

原文摘要 · Abstract (English)

Large Language Models (LLM)-based Multi-Agent Systems (MASs) have emerged as a new paradigm in software system design, increasingly demonstrating strong reasoning and collaboration capabilities. As these systems become more complex and autonomous, effective failure management is essential to ensure reliability and availability. However, existing approaches often rely on per-trace reasoning, which leads to low efficiency, and neglect historical failure patterns, limiting diagnostic accuracy. In this paper, we conduct a preliminary empirical study to demonstrate the necessity, potential, and challenges of leveraging historical failure patterns to enhance failure management in MASs. Building on this insight, we propose \textbf{EAGER}, an efficient failure management framework for multi-agent systems based on reasoning trace representation. EAGER employs unsupervised reasoning-scoped contrastive learning to encode both intra-agent reasoning and inter-agent coordination, enabling real-time step-wise failure detection, diagnosis, and reflexive mitigation guided by historical failure knowledge. Preliminary evaluations on three open-source MASs demonstrate the effectiveness of EAGER and highlight promising directions for future research in reliable multi-agent system operations.

多智能体故障管理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。