arXiv:2410.16251cs.CL2024-10ICLR被引 41

提出新基准,验证知识编辑能否真正消除大模型幻觉。

Can Knowledge Editing Really Correct Hallucinations?

  • 构建包含6000+幻觉的跨领域评测集,确保模型先产生幻觉再测试修复效果。
  • 在5个维度上评估多种编辑方法,发现现有技术仍存局限。
  • 为知识编辑研究提供可复现的评测框架,适合关注模型可信度的开发者。

大型语言模型虽具备强大能力,但仍存在生成非事实内容的幻觉问题。知识编辑作为无需重新训练即可修正模型错误知识的新范式,其有效性尚待验证。现有评测数据集未保证模型在编辑前确实产生幻觉,导致无法准确评估不同编辑方法的纠错能力。为此,我们提出 HalluEditBench,系统性地评估知识编辑在真实幻觉修正中的表现。该基准涵盖9个领域、26个主题,包含超过6000个经过严格验证的幻觉样本,并从有效性、泛化性、可迁移性、局部性和鲁棒性五个维度综合评估多种编辑方法。实验揭示了当前方法的优势与不足,为未来研究提供了关键洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) suffer from hallucinations, referring to the non-factual information in generated content, despite their superior capacities across tasks. Meanwhile, knowledge editing has been developed as a new popular paradigm to correct erroneous factual knowledge encoded in LLMs with the advantage of avoiding retraining from scratch. However, a common issue of existing evaluation datasets for knowledge editing is that they do not ensure that LLMs actually generate hallucinated answers to the evaluation questions before editing. When LLMs are evaluated on such datasets after being edited by different techniques, it is hard to directly adopt the performance to assess the effectiveness of different knowledge editing methods in correcting hallucinations. Thus, the fundamental question remains insufficiently validated: Can knowledge editing really correct hallucinations in LLMs? We proposed HalluEditBench to holistically benchmark knowledge editing methods in correcting real-world hallucinations. First, we rigorously construct a massive hallucination dataset with 9 domains, 26 topics and more than 6,000 hallucinations. Then, we assess the performance of knowledge editing methods in a holistic way on five dimensions including Efficacy, Generalization, Portability, Locality, and Robustness. Through HalluEditBench, we have provided new insights into the potentials and limitations of different knowledge editing methods in correcting hallucinations, which could inspire future improvements and facilitate progress in the field of knowledge editing.

知识编辑幻觉纠正评测基准大模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。