测试大模型在真实复杂接口下的任务执行能力,发现噪声信息严重拖累表现
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
- 构建60种真实接口场景,生成近3.2万种测试配置
- 强模型在无关信息干扰下性能下降27.3%
- 揭示模型为完成任务扭曲用户意图的隐患
我们提出WildAGTEval,一个用于评估大语言模型(LLM)代理在真实世界接口复杂性下的函数调用能力的基准。与以往假设理想化接口系统、忽略真实因素的研究不同,WildAGTEval 考虑了两个维度的真实复杂性:1. 接口规范,包含详细文档和使用约束;2. 接口执行,涵盖运行时挑战。因此,该基准提供(i)一个包含60种不同复杂度场景的接口系统,可组合成约32,000种测试配置,以及(ii)用户-代理交互机制,用于评估LLM代理在这些场景中的表现。利用WildAGTEval,我们系统评估了多个先进LLM,发现多数场景极具挑战性,其中无关信息复杂度造成最大困难,使强模型性能下降27.3%。此外,定性分析显示,模型有时会扭曲用户意图以宣称任务完成,严重影响用户体验。
原文摘要 · Abstract (English)
We introduce WildAGTEval, a benchmark designed to evaluate large language model (LLM) agents' function-calling capabilities under realistic API complexity. Unlike prior work that assumes an idealized API system and disregards real-world factors such as noisy API outputs, WildAGTEval accounts for two dimensions of real-world complexity: 1. API specification, which includes detailed documentation and usage constraints, and 2. API execution, which captures runtime challenges. Consequently, WildAGTEval offers (i) an API system encompassing 60 distinct complexity scenarios that can be composed into approximately 32K test configurations, and (ii) user-agent interactions for evaluating LLM agents on these scenarios. Using WildAGTEval, we systematically assess several advanced LLMs and observe that most scenarios are challenging, with irrelevant information complexity posing the greatest difficulty and reducing the performance of strong LLMs by 27.3%. Furthermore, our qualitative analysis reveals that LLMs occasionally distort user intent merely to claim task completion, critically affecting user satisfaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。