arXiv:2504.18373cs.CL2025-04EMNLP被引 2

构建智能助手多智能体评估基准,推动真实场景下系统能力测试

Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant

  • 扩展SLURP数据集,加入模拟服务与任务执行链路
  • 实现从理解到响应的全流程端到端评估,验证当前框架仍不成熟
  • 适合研究多智能体系统、智能助手可靠性的学者和开发者

近年来,基于大语言模型(LLMs)的多智能体框架发展迅速。然而,针对其性能评估的专用基准数据集仍属空白。为此,我们提出Auto-SLURP,一个面向智能个人助手场景的多智能体框架评估基准。该数据集在原始SLURP数据集基础上,通过重新标注并集成模拟服务器与外部服务,实现了涵盖语言理解、任务执行与响应生成的全流程端到端评估。实验表明,Auto-SLURP对当前最先进框架构成显著挑战,凸显真正可靠、智能的多智能体个人助手尚处于探索阶段。相关数据与代码已开源:https://github.com/lorashen/Auto-SLURP/

原文摘要 · Abstract (English)

In recent years, multi-agent frameworks powered by large language models (LLMs) have advanced rapidly. Despite this progress, there is still a notable absence of benchmark datasets specifically tailored to evaluate their performance. To bridge this gap, we introduce Auto-SLURP, a benchmark dataset aimed at evaluating LLM-based multi-agent frameworks in the context of intelligent personal assistants. Auto-SLURP extends the original SLURP dataset -- initially developed for natural language understanding tasks -- by relabeling the data and integrating simulated servers and external services. This enhancement enables a comprehensive end-to-end evaluation pipeline, covering language understanding, task execution, and response generation. Our experiments demonstrate that Auto-SLURP presents a significant challenge for current state-of-the-art frameworks, highlighting that truly reliable and intelligent multi-agent personal assistants remain a work in progress. The dataset and related code are available at https://github.com/lorashen/Auto-SLURP/.

多智能体智能助手评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。