诊断大模型多智能体系统工具调用失败原因,提升企业自动化可靠性
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
- 构建12类错误分类体系,系统分析工具初始化、参数处理等环节故障
- 测试1980个实例发现小模型工具初始化失败率高,qwen2.5:32b表现媲美GPT-4
- 中等规模模型在普通硬件上达96.6%成功率,适合资源有限企业部署
基于大语言模型的多智能体系统正推动企业自动化变革,但工具使用可靠性的系统性评估方法仍不成熟。本文提出一个综合诊断框架,利用大数据分析评估智能体系统的流程可靠性,满足中小企业在隐私敏感环境中的部署需求。该框架包含12类错误分类,涵盖工具初始化、参数处理、执行与结果解析等环节。通过对1,980个确定性测试实例的系统评估,覆盖开源模型(Qwen2.5系列、Functionary)和专有模型(GPT-4、Claude 3.5/3.7),在多种边缘硬件配置下进行验证,识别出生产部署的可操作可靠性阈值。分析显示,中小模型的流程可靠性瓶颈主要来自工具初始化失败;而qwen2.5:32b实现零错误表现,与GPT-4.1相当。中等规模模型(qwen2.5:14b)在通用硬件上达到96.6%成功率、7.3秒延迟,为资源受限组织提供成本效益高的智能体部署方案。本工作建立了工具增强型多智能体系统可靠性评估的基础架构。
原文摘要 · Abstract (English)
Multi-agent systems powered by large language models (LLMs) are transforming enterprise automation, yet systematic evaluation methodologies for assessing tool-use reliability remain underdeveloped. We introduce a comprehensive diagnostic framework that leverages big data analytics to evaluate procedural reliability in intelligent agent systems, addressing critical needs for SME-centric deployment in privacy-sensitive environments. Our approach features a 12-category error taxonomy capturing failure modes across tool initialization, parameter handling, execution, and result interpretation. Through systematic evaluation of 1,980 deterministic test instances spanning both open-weight models (Qwen2.5 series, Functionary) and proprietary alternatives (GPT-4, Claude 3.5/3.7) across diverse edge hardware configurations, we identify actionable reliability thresholds for production deployment. Our analysis reveals that procedural reliability, particularly tool initialization failures, constitutes the primary bottleneck for smaller models, while qwen2.5:32b achieves flawless performance matching GPT-4.1. The framework demonstrates that mid-sized models (qwen2.5:14b) offer practical accuracy-efficiency trade-offs on commodity hardware (96.6\% success rate, 7.3 s latency), enabling cost-effective intelligent agent deployment for resource-constrained organizations. This work establishes foundational infrastructure for systematic reliability evaluation of tool-augmented multi-agent AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。