arXiv:2411.00034cs.CLcs.AI2024-11

用自动化方法检测荷兰语客服聊天机器人答错的情况。

Is Our Chatbot Telling Lies? Assessing Correctness of an LLM-based Dutch Support Chatbot

  • 基于真实客服决策定义正确性,构建自动化评估框架。
  • 对错误回答的识别率达55%,可实时发现潜在误导。
  • 适合关注本地化语言与问答类型优化的AI客服开发者。

企业通过在线聊天和聊天机器人提升客户忠诚度。AFAS是一家荷兰公司,旨在利用大语言模型(LLMs)在极少人工干预下回应客户咨询。然而,何为正确回答尚不明确,尤其在荷兰语场景中。加之训练数据有限,如何即时判断大模型生成的回答是否正确成为挑战。本研究首次依据AFAS客服团队的实际决策标准,定义回答正确性,并结合自然语言生成与自动评分文献,实现对客服决策的自动化模拟。我们测试了需二元回答(如:能否手动调整税率?)或指令类问题(如:如何手动调整税率?),结果表明该方法能在55%的情况下识别出错误信息。研究展示了自动检测聊天机器人可能产生错误或误导性回答的潜力,贡献包括:(1) 提出正确性的定义与评估指标;(2) 针对区域语言和问题类型提出改进建议。

原文摘要 · Abstract (English)

Companies support their customers using live chats and chatbots to gain their loyalty. AFAS is a Dutch company aiming to leverage the opportunity large language models (LLMs) offer to answer customer queries with minimal to no input from its customer support team. Adding to its complexity, it is unclear what makes a response correct, and that too in Dutch. Further, with minimal data available for training, the challenge is to identify whether an answer generated by a large language model is correct and do it on the fly. This study is the first to define the correctness of a response based on how the support team at AFAS makes decisions. It leverages literature on natural language generation and automated answer grading systems to automate the decision-making of the customer support team. We investigated questions requiring a binary response (e.g., Would it be possible to adjust tax rates manually?) or instructions (e.g., How would I adjust tax rate manually?) to test how close our automated approach reaches support rating. Our approach can identify wrong messages in 55\% of the cases. This work demonstrates the potential for automatically assessing when our chatbot may provide incorrect or misleading answers. Specifically, we contribute (1) a definition and metrics for assessing correctness, and (2) suggestions to improve correctness with respect to regional language and question type.

大模型评估客服机器人荷兰语自动评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。