用评分制量化生成内容的幻觉程度,提升事实一致性评估效率
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation
- 设计评分量表,用大模型自动判断生成内容与源文本的一致性
- GPT-4在4个旅行数据集上准确识别事实错误,合成数据可提升NLI模型表现
- 适合需要高可信度生成的工业应用,尤其关注成本与延迟的部署者
幻觉是自然语言生成中的热点问题。在真实场景中,不忠实的内容会降低数据质量或损害用户信任,因此在生产环境中使用NLG前必须进行事实核查,而人工核查成本高昂。本文研究了引导式NLG中的自动化忠实度评估,开发了一套评分量表,并利用大语言模型(LLMs)对生成内容进行可量化打分。我们对比了主流LLM与广泛使用的自然语言推理(NLI)模型在评分质量与敏感性上的表现。此外,我们提出了合成非忠实数据的方法及量化幻觉比例的启发式策略。在4个旅行领域的真实行业数据集上的实验表明,GPT-4能准确判断源文本与生成内容的事实一致性,并提供合理解释。同时发现,在合成数据上微调NLI模型可提升性能。最后,我们分析了该系统部署时的延迟与成本。
原文摘要 · Abstract (English)
Hallucination has been a popular topic in natural language generation (NLG). In real-world applications, unfaithful content can result in poor data quality or loss of trust from end users. Thus, it is crucial to fact-check before adopting NLG for production usage, which can be expensive if done manually. In this paper, we investigate automated faithfulness evaluation in guided NLG. We developed a rubric template and used large language models (LLMs) to score the generation on quantifiable scales. We compared popular LLMs as well as widely adopted natural language inference (NLI) models in scoring quality and sensitivity. In addition, we developed methods for the generation of synthetic unfaithful data, as well as heuristics to quantify the percentage of hallucination. Our results on 4 travel-domain industry dataset show that GPT-4 can provide accurate judgement and explanation of whether a source and a generation are factually consistent. Furthermore, we found that tuning NLI models on synthetic data can improve performance. Lastly, we present insights on the latency and cost of deploying such a system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。