arXiv:2501.00269cs.CL2025-01综述被引 22

综述大模型幻觉评估中的忠实度度量方法,指出用大模型评估忠实度最接近人工判断。

A review of faithfulness metrics for hallucination assessment in Large Language Models

  • 以大模型自身作为忠实度评估器,与人工判断相关性最高。
  • 检索增强生成和提示框架能有效提升生成内容的忠实度。
  • 适合关注大模型可靠性与可信度的研究者阅读。

本文综述了在开放问答、问题回答和机器翻译任务中忠实度评估的方法。研究发现,使用大语言模型作为忠实度评估指标时,其与人工判断的相关性最高。文章还讨论了其他缓解幻觉的方法,指出检索增强生成(RAG)和提示框架均与更高的忠实度相关,同时提供了其他缓解建议。忠实度研究对大模型的广泛应用至关重要,因为不忠实的输出可能在多个领域带来重大风险。此外,开放生成任务的评估比常见的多选基准更能全面反映大模型性能,有助于提升对大模型的信任度。

原文摘要 · Abstract (English)

This review examines the means with which faithfulness has been evaluated across open-ended summarization, question-answering and machine translation tasks. We find that the use of LLMs as a faithfulness evaluator is commonly the metric that is most highly correlated with human judgement. The means with which other studies have mitigated hallucinations is discussed, with both retrieval augmented generation (RAG) and prompting framework approaches having been linked with superior faithfulness, whilst other recommendations for mitigation are provided. Research into faithfulness is integral to the continued widespread use of LLMs, as unfaithful responses can pose major risks to many areas whereby LLMs would otherwise be suitable. Furthermore, evaluating open-ended generation provides a more comprehensive measure of LLM performance than commonly used multiple-choice benchmarking, which can help in advancing the trust that can be placed within LLMs.

大模型幻觉评估忠实度综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。