用大模型自动评测车载对话系统的事实正确性,准确率超90%。
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models
- 基于大模型设计自动评测框架,结合提示工程与多角色协同减少幻觉。
- 对车载问答系统测试中,GPT-4在输入输出提示下达到90%以上正确率。
- 适合需高可靠性对话系统验证的自动驾驶与智能座舱研发人员。
车载对话系统有望提升驾乘体验。现代系统依赖大语言模型(LLM),易产生幻觉——即不准确、虚构的错误信息。本文提出一种基于LLM的自动化事实基准评测方法,通过五种基于LLM的实现方式,结合集成技术与多样化角色设定以提高一致性并减少幻觉。我们用该方法评估了CarExpert——一个基于检索增强的车载问答系统——在车辆手册事实正确性方面的表现。构建了专用于车载场景的新数据集,并与专家评估进行对比。结果显示,使用GPT-4配合输入输出提示的方法,在事实正确性上与专家评估的吻合率达90%以上,且平均响应时间仅为4.5秒,是效率最高的方案。研究表明,基于大模型的测试可有效验证对话系统的真实可靠性。
原文摘要 · Abstract (English)
In-car conversational systems bring the promise to improve the in-vehicle user experience. Modern conversational systems are based on Large Language Models (LLMs), which makes them prone to errors such as hallucinations, i.e., inaccurate, fictitious, and therefore factually incorrect information. In this paper, we present an LLM-based methodology for the automatic factual benchmarking of in-car conversational systems. We instantiate our methodology with five LLM-based methods, leveraging ensembling techniques and diverse personae to enhance agreement and minimize hallucinations. We use our methodology to evaluate CarExpert, an in-car retrieval-augmented conversational question answering system, with respect to the factual correctness to a vehicle's manual. We produced a novel dataset specifically created for the in-car domain, and tested our methodology against an expert evaluation. Our results show that the combination of GPT-4 with the Input Output Prompting achieves over 90 per cent factual correctness agreement rate with expert evaluations, other than being the most efficient approach yielding an average response time of 4.5s. Our findings suggest that LLM-based testing constitutes a viable approach for the validation of conversational systems regarding their factual correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。