用知识图谱增强文本检索,让大模型回答更准确。
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
- 将知识图谱与语义检索结合,融合结构化与非结构化数据。
- 在意大利美食数据集上答相关性达0.94,上下文精确率达0.99。
- 适合需要高准确率和上下文一致性的问答系统开发者。
检索增强生成(RAG)系统通过利用海量语料生成有依据且上下文相关的回应,显著降低大语言模型的幻觉问题。尽管取得进展,现有系统在处理大规模数据时仍难以高效检索并保持全面的上下文理解。本文提出SKETCH,一种新型方法,通过整合语义文本检索与知识图谱,融合结构化与非结构化数据,实现更全面的语义理解。在四个不同数据集(QuALITY、QASPER、NarrativeQA、Italian Cuisine)上的评估显示,SKETCH在关键RAGAS指标如答案相关性、忠实性、上下文精确率和上下文召回率上均优于基线方法。尤其在Italian Cuisine数据集上,答案相关性达0.94,上下文精确率达0.99,为所有评测指标最高表现。结果表明SKETCH能生成更准确、上下文一致的回应,为未来检索系统树立新基准。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) systems have become pivotal in leveraging vast corpora to generate informed and contextually relevant responses, notably reducing hallucinations in Large Language Models. Despite significant advancements, these systems struggle to efficiently process and retrieve information from large datasets while maintaining a comprehensive understanding of the context. This paper introduces SKETCH, a novel methodology that enhances the RAG retrieval process by integrating semantic text retrieval with knowledge graphs, thereby merging structured and unstructured data for a more holistic comprehension. SKETCH, demonstrates substantial improvements in retrieval performance and maintains superior context integrity compared to traditional methods. Evaluated across four diverse datasets: QuALITY, QASPER, NarrativeQA, and Italian Cuisine-SKETCH consistently outperforms baseline approaches on key RAGAS metrics such as answer_relevancy, faithfulness, context_precision and context_recall. Notably, on the Italian Cuisine dataset, SKETCH achieved an answer relevancy of 0.94 and a context precision of 0.99, representing the highest performance across all evaluated metrics. These results highlight SKETCH's capability in delivering more accurate and contextually relevant responses, setting new benchmarks for future retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。