arXiv:2501.12851cs.CL2025-01被引 59

新基准ACEBench评估大模型用工具能力,覆盖真实对话场景。

ACEBench: Who Wins the Match Point in Tool Usage?

  • 分三类数据:基础、模糊指令、多智能体对话,模拟真实使用
  • 测试多个大模型在复杂任务中调用工具的准确率与错误类型
  • 适合研究工具增强型模型或评估系统可靠性的团队

大型语言模型(LLMs)在决策与推理方面展现出巨大潜力,尤其在整合各类工具解决复杂问题时。然而,现有评估大模型工具使用的基准存在诸多局限:(1) 评估场景有限,常缺乏对真实多轮对话情境的测试;(2) 评估维度狭窄,对工具使用过程的分析不够深入;(3) 依赖大模型自身或真实API执行进行评估,带来显著开销。为此,我们提出ACEBench,一个全面的大模型工具使用评估基准。该基准依据评估方法将数据分为三类:Normal(基础场景)、Special(指令模糊或不完整)、Agent(多智能体交互,模拟真实多轮对话)。我们在ACEBench上展开大量实验,深入分析多种大模型的表现,并对不同数据类型下的错误原因进行更细致的剖析。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs' tool usage face several limitations: (1) limited evaluation scenarios, often lacking assessments in real multi-turn dialogue contexts; (2) narrow evaluation dimensions, with insufficient detailed assessments of how LLMs use tools; and (3) reliance on LLMs or real API executions for evaluation, which introduces significant overhead. To address these challenges, we introduce ACEBench, a comprehensive benchmark for assessing tool usage in LLMs. ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent. "Normal" evaluates tool usage in basic scenarios; "Special" evaluates tool usage in situations with ambiguous or incomplete instructions; "Agent" evaluates tool usage through multi-agent interactions to simulate real-world, multi-turn dialogues. We conducted extensive experiments using ACEBench, analyzing various LLMs in-depth and providing a more granular examination of error causes across different data types.

大模型评估工具使用多轮对话基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。