arXiv:2411.09972cs.CLcs.AI2024-11被引 11

用大模型模拟真实用户,动态评估对话系统表现

Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems

  • 用大模型生成有上下文感知的虚拟用户
  • 改进提示词后,用户行为多样性和任务完成率提升
  • 提供自动化评估框架,适合研究对话系统效能

传统离线数据集缺乏上下文感知,难以有效评估任务导向型对话(TOD)系统。相比之下,具备上下文意识的用户代理能更好模拟人类对话的多样性和不可预测性,是更优的评估方式。本文利用大语言模型(LLM)构建用户代理,通过上下文示例引导并追踪用户目标状态。实验表明,优化提示词可显著提升用户代理在多样性与任务完成度上的表现。同时,本文提出在该动态框架下对TOD模型进行自动评估的方法。

原文摘要 · Abstract (English)

Traditionally, offline datasets have been used to evaluate task-oriented dialogue (TOD) models. These datasets lack context awareness, making them suboptimal benchmarks for conversational systems. In contrast, user-agents, which are context-aware, can simulate the variability and unpredictability of human conversations, making them better alternatives as evaluators. Prior research has utilized large language models (LLMs) to develop user-agents. Our work builds upon this by using LLMs to create user-agents for the evaluation of TOD systems. This involves prompting an LLM, using in-context examples as guidance, and tracking the user-goal state. Our evaluation of diversity and task completion metrics for the user-agents shows improved performance with the use of better prompts. Additionally, we propose methodologies for the automatic evaluation of TOD models within this dynamic framework.

对话系统大模型评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。