为大模型驱动的智能体设计专用架构评估方法
AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
- 提出针对大模型智能体特性的新型架构评估方法
- 构建了面向智能体的通用场景目录,指导实际评估
- 在真实税务助手系统上验证方法有效性
基础模型(FMs)的出现推动了高度自主智能体的发展,开启了跨领域的新应用机遇。由于智能体具有复合架构、自主非确定性行为及持续演化的独特特性,其架构设计对智能体质量属性有显著影响,因此架构评估尤为重要。然而,传统评估方法难以满足此类智能体的评估需求。为此,本文提出 AgentArcEval,一种专为基于基础模型的智能体架构及其评估复杂性而设计的新型评估方法。此外,我们构建了一个面向智能体的通用场景目录,可作为生成具体评估场景的指南。通过一个真实税务助手系统 Luna 的案例研究,展示了 AgentArcEval 及其场景目录的有效性。
原文摘要 · Abstract (English)
The emergence of foundation models (FMs) has enabled the development of highly capable and autonomous agents, unlocking new application opportunities across a wide range of domains. Evaluating the architecture of agents is particularly important as the architectural decisions significantly impact the quality attributes of agents given their unique characteristics, including compound architecture, autonomous and non-deterministic behaviour, and continuous evolution. However, these traditional methods fall short in addressing the evaluation needs of agent architecture due to the unique characteristics of these agents. Therefore, in this paper, we present AgentArcEval, a novel agent architecture evaluation method designed specially to address the complexities of FM-based agent architecture and its evaluation. Moreover, we present a catalogue of agent-specific general scenarios, which serves as a guide for generating concrete scenarios to design and evaluate the agent architecture. We demonstrate the usefulness of AgentArcEval and the catalogue through a case study on the architecture evaluation of a real-world tax copilot, named Luna.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。