剖析低代码平台中智能体工作流的故障生命周期,揭示其失效机制与修复方法。
Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
- 通过实证研究分析故障传播路径与根因
- 构建包含307个真实故障案例的数据集
- 为可靠设计与故障修复提供可操作指南
基于低代码编排平台构建的智能体工作流虽能加速多智能体系统开发,但引入了新型且尚不明确的故障模式,影响系统的可靠性与可维护性。与传统软件不同,此类工作流中的故障常通过自然语言交互、工具调用和动态控制逻辑在异构节点间传播,导致故障归因与修复极为困难。本文从故障生命周期视角出发,对两类代表性平台上的智能体工作流进行实证研究,旨在刻画故障表现、识别根本原因并分析修复策略。我们构建了名为AgentFail的数据集,包含307个真实世界故障案例。基于此数据集,分析了不同故障根因及工作流节点的故障模式、根本原因与修复难度。研究揭示了智能体工作流中的关键故障机制,并提出了可操作的可靠性修复建议与实际设计指导。
原文摘要 · Abstract (English)
Agentic workflows built on low-code orchestration platforms enable rapid development of multi-agent systems, but they also introduce new and poorly understood failure modes that hinder reliability and maintainability. Unlike traditional software systems, failures in agentic workflows often propagate across heterogeneous nodes through natural-language interactions, tool invocations, and dynamic control logic, making failure attribution and repair particularly challenging. In this paper, we present an empirical study of platform-orchestrated agentic workflows from a failure lifecycle perspective, with the goal of characterizing failure manifestations, identifying underlying root causes, and examining corresponding repair strategies. We present AgentFail, a dataset of 307 real-world failure cases collected from two representative agentic workflow platforms. Based on this dataset, we analyze failure patterns, root causes, and repair difficulty for various failure root causes and nodes in the workflow. Our findings reveal key failure mechanisms in agentic workflows and provide actionable guidelines for reliable failure repair, and real-world agentic workflow design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。