arXiv:2411.14572cs.LGcs.CL2024-11被引 23

用语言模型的表征提升检索增强生成的知识可信度

Towards Knowledge Checking in Retrieval-augmented Generation: A Representation Perspective

  • 基于大模型内部表征设计知识筛选分类器
  • 在噪声数据库下仍显著提升RAG性能
  • 适合关注生成可靠性与知识整合的研究者

检索增强生成(RAG)系统在提升大语言模型(LLM)性能方面展现出潜力。然而,这些系统在有效融合外部知识与模型内部知识时面临挑战,常导致误导性或无益信息。本文对RAG系统中的知识检查进行系统性研究,全面分析了LLM表征行为,并证明使用表征进行知识检查的重要性。基于此发现,我们进一步开发了基于表征的知识过滤分类器。实验表明,即使在噪声知识数据库条件下,也能显著提升RAG性能。本研究为利用LLM表征增强RAG系统的可靠性与有效性提供了新视角。

原文摘要 · Abstract (English)

Retrieval-Augmented Generation (RAG) systems have shown promise in enhancing the performance of Large Language Models (LLMs). However, these systems face challenges in effectively integrating external knowledge with the LLM's internal knowledge, often leading to issues with misleading or unhelpful information. This work aims to provide a systematic study on knowledge checking in RAG systems. We conduct a comprehensive analysis of LLM representation behaviors and demonstrate the significance of using representations in knowledge checking. Motivated by the findings, we further develop representation-based classifiers for knowledge filtering. We show substantial improvements in RAG performance, even when dealing with noisy knowledge databases. Our study provides new insights into leveraging LLM representations for enhancing the reliability and effectiveness of RAG systems.

RAG知识检查表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。