用教师模型评估聊天机器人是否达成用户目标,提升评估效率与可解释性。
Mind the Goal: Data-Efficient Goal-Oriented Evaluation of Conversational Agents and Chatbots using Teacher Models
- 基于教师大模型和思维令牌,按用户目标分割对话并评估成功率
- 在企业场景中将目标达成率从63%提升至79%,六个月内持续优化
- 提供失败原因分类体系,帮助诊断系统缺陷并指导改进
多轮对话评估仍具挑战性,现有方法多在回合层面评价,未能判断用户整体目标是否达成。本文提出面向多智能体系统的整体目标评估框架,引入目标成功率达(GSR)衡量目标完成比例,并建立失败根因(RCOF)分类体系以识别多智能体聊天机器人失败原因。该方法按用户目标分割对话,综合相关回合评估。采用基于教师大模型的评估系统,由领域专家定义目标与质量标准,指导大模型生成带有“思考令牌”的可解释推理过程,实现可解释、数据高效的评估。在企业场景中,对零到一构建的员工对话系统AIDA应用此框架,目标成功率在六个月内从63%提升至79%。该框架通用性强,通过失败点分析提供行动洞察,诊断整体成功率,识别关键失败模式,推动系统优化。
原文摘要 · Abstract (English)
Evaluating the quality of multi-turn chatbot interactions remains challenging, as most existing methods assess interactions at the turn level without addressing whether a user's overarching goal was fulfilled. A ``goal'' here refers to an information need or task, such as asking for policy information or applying for leave. We propose a comprehensive framework for goal-oriented evaluation of multi-agent systems (MAS), introducing the \textbf{Goal Success Rate (GSR)} to measure the percentage of fulfilled goals, and a \textbf{Root Cause of Failure (RCOF)} taxonomy to identify reasons for failure in multi-agent chatbots. Our method segments conversations by user goals and evaluates success using all relevant turns. We present a model-based evaluation system combining teacher LLMs, where domain experts define goals, set quality standards serving as a guidance for the LLMs. The LLMs use ``thinking tokens'' to produce interpretable rationales, enabling \textit{explainable}, \textit{data-efficient} evaluations. In an enterprise setting, we apply our framework to evaluate AIDA, a zero-to-one employee conversational agent system built as a ground-up multi-agent conversational agent, and observe GSR improvement from 63\% to 79\% over six months since its inception. Our framework is generic and offers actionable insights through a detailed defect taxonomy based on analysis of failure points in multi-agent chatbots, diagnosing overall success, identifying key failure modes, and informing system improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。