arXiv:2409.15934cs.CLcs.AI2024-09被引 19

用自动化测试评估大模型对话代理的全流程能力

Automated test generation to evaluate tool-augmented LLMs as conversational AI agents

  • 用中间图约束生成内容,避免幻觉并覆盖多种对话路径
  • 在客户支持场景中发现大模型能单次调用但难处理完整对话
  • 方法通用,适用于多领域对话智能体评测

工具增强的大模型是构建可进行真实对话、遵循流程并调用合适函数的AI代理的有前景方法。然而,由于对话可能形态多样,评估极具挑战性,现有数据集仅关注单次交互和函数调用。本文提出一个测试生成流水线,用于评估大模型作为对话式AI代理的表现。我们的框架利用大模型生成基于用户定义流程的多样化测试,通过中间图限制测试生成器的幻觉倾向,并确保对可能对话路径的高覆盖率。此外,我们构建了ALMITA——一个手动标注的客户支持场景评估数据集,并用于测试现有大模型。结果表明,尽管工具增强的大模型在单次交互中表现良好,但在处理完整对话时经常遇到困难。虽然研究聚焦于客户支持,但该方法具有普适性,可用于不同领域的智能体评估。

原文摘要 · Abstract (English)

Tool-augmented LLMs are a promising approach to create AI agents that can have realistic conversations, follow procedures, and call appropriate functions. However, evaluating them is challenging due to the diversity of possible conversations, and existing datasets focus only on single interactions and function-calling. We present a test generation pipeline to evaluate LLMs as conversational AI agents. Our framework uses LLMs to generate diverse tests grounded on user-defined procedures. For that, we use intermediate graphs to limit the LLM test generator's tendency to hallucinate content that is not grounded on input procedures, and enforces high coverage of the possible conversations. Additionally, we put forward ALMITA, a manually curated dataset for evaluating AI agents in customer support, and use it to evaluate existing LLMs. Our results show that while tool-augmented LLMs perform well in single interactions, they often struggle to handle complete conversations. While our focus is on customer support, our method is general and capable of AI agents for different domains.

对话系统自动化测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。