系统梳理智能体运维框架,解决大模型系统不稳定难题
A Survey on AgentOps: Categorization, Challenges, and Future Directions
- 按异常来源分为智能体内部与跨智能体两类
- 提出四阶段运维框架:监控、检测、根因分析、修复
- 为落地部署的大模型智能体提供可操作的维护方案
随着大语言模型推理能力的不断提升,基于大语言模型的智能体系统在灵活性和可解释性上优于传统系统,受到广泛关注。然而,尽管智能体系统在学术界和工业界均有广泛应用,其仍频繁出现异常,导致系统不稳定与不安全,制约进一步发展。因此,亟需一套全面、系统的智能体运维方法。但当前关于智能体运维的研究仍较为匮乏。为此,本文开展智能体运维(AgentOps)综述,旨在建立清晰的领域框架,定义关键挑战并推动未来发展。具体而言,本文首先系统定义智能体系统中的异常,将其分类为内部异常与跨智能体异常;随后提出一个全新的、全面的智能体系统运维框架——AgentOps,详细阐述其四个核心阶段:监控、异常检测、根因分析与恢复。
原文摘要 · Abstract (English)
As the reasoning capabilities of Large Language Models (LLMs) continue to advance, LLM-based agent systems offer advantages in flexibility and interpretability over traditional systems, garnering increasing attention. However, despite the widespread research interest and industrial application of agent systems, these systems, like their traditional counterparts, frequently encounter anomalies. These anomalies lead to instability and insecurity, hindering their further development. Therefore, a comprehensive and systematic approach to the operation and maintenance of agent systems is urgently needed. Unfortunately, current research on the operations of agent systems is sparse. To address this gap, we have undertaken a survey on agent system operations with the aim of establishing a clear framework for the field, defining the challenges, and facilitating further development. Specifically, this paper begins by systematically defining anomalies within agent systems, categorizing them into intra-agent anomalies and inter-agent anomalies. Next, we introduce a novel and comprehensive operational framework for agent systems, dubbed Agent System Operations (AgentOps). We provide detailed definitions and explanations of its four key stages: monitoring, anomaly detection, root cause analysis, and resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。