arXiv:2409.16658cs.CL2024-09EMNLP被引 1

大模型生成幻觉文本时概率分布可区分,据此可减少幻觉。

Pre-trained Language Models Return Distinguishable Probability Distributions to Unfaithfully Hallucinated Texts

  • 利用生成概率与不确定性分布差异识别幻觉文本
  • 88%-98%情况下能显著区分幻觉与真实文本
  • 新训练算法提升忠实度,保持文本质量

本文表明,预训练语言模型对不忠实的幻觉文本会产生可区分的生成概率与不确定性分布,无论其规模和结构如何。通过对6个数据集上的24个模型进行分析,发现88%-98%的情况返回统计上显著可区分的概率与不确定性分布。基于这一普遍现象,我们提出一种减少幻觉的训练算法。该算法在保持良好通用文本质量的同时,优于其他基线方法,在忠实度指标上表现更优。

原文摘要 · Abstract (English)

In this work, we show the pre-trained language models return distinguishable generation probability and uncertainty distribution to unfaithfully hallucinated texts, regardless of their size and structure. By examining 24 models on 6 data sets, we find out that 88-98% of cases return statistically significantly distinguishable generation probability and uncertainty distributions. Using this general phenomenon, we showcase a hallucination-reducing training algorithm. Our algorithm outperforms other baselines by achieving higher faithfulness metrics while maintaining sound general text quality measures.

幻觉检测语言模型概率分布可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。