arXiv:2411.07870cs.CLcs.AI2024-11被引 1

用知识库和双解码器让大模型生成更可信的领域内容

Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders

  • 用知识三元组后处理纠正生成中的幻觉
  • 双解码器融合检索信息,提升生成准确性
  • 适合需要高可信度文本生成的垂直领域

尽管大语言模型展现出强大的内容生成能力,但其应用受限于内容的领域依赖性。生成内容的正确性和事实依据需基于经验证的上下文,如检索增强生成(RAG)结果。将大模型适配至定制化领域时,常出现回复不完整或新增内容未经验证,甚至产生幻觉。现有幻觉检测研究多关注评估指标,难以适应动态领域,且易受越狱攻击影响。本文提出:1)利用RAG上下文中知识三元组的后处理算法以纠正幻觉;2)一种双解码器模型,将RAG上下文融入生成过程,实现更可靠的文本生成。

原文摘要 · Abstract (English)

Although people are impressed by the content generation skills of large language models, the use of LLMs, such as ChatGPT, is limited by the domain grounding of the content. The correctness and groundedness of the generated content need to be based on a verified context, such as results from Retrieval-Augmented Generation (RAG). One important issue when adapting LLMs to a customized domain is that the generated responses are often incomplete, or the additions are not verified and may even be hallucinated. Prior studies on hallucination detection have focused on evaluation metrics, which are not easily adaptable to dynamic domains and can be vulnerable to attacks like jail-breaking. In this work, we propose 1) a post-processing algorithm that leverages knowledge triplets in RAG context to correct hallucinations and 2) a dual-decoder model that fuses RAG context to guide the generation process.

大模型知识库幻觉纠正RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。