用大模型+检索增强,让波斯语大学问答更准更快
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval
- 结合网页数据与提示工程构建检索增强系统
- 在自建基准上准确率显著提升,响应更相关
- 适合需要精准学术信息检索的教育类应用
本文提出一种基于检索增强生成(RAG)与大语言模型(LLMs)的创新方法,用于提升大学相关问题问答系统的性能。通过系统提取高校官网数据,并采用先进提示工程,生成准确且上下文相关的回答。我们构建了全面的大学问答基准 UniversityQuestionBench(UQB),基于领域内常见指标评估系统表现,涵盖准确率与可靠性,在多种真实场景下进行测试。实验结果表明,生成回答的精确度和相关性显著提高,用户获取有效信息的时间大幅缩短。本研究展示了RAG与LLM在学术数据检索中的新应用,依托精心设计的基准,为未来该领域研究奠定基础。
原文摘要 · Abstract (English)
This paper introduces an innovative approach using Retrieval-Augmented Generation (RAG) pipelines with Large Language Models (LLMs) to enhance information retrieval and query response systems for university-related question answering. By systematically extracting data from the university official webpage and employing advanced prompt engineering techniques, we generate accurate, contextually relevant responses to user queries. We developed a comprehensive university benchmark, UniversityQuestionBench (UQB), to rigorously evaluate our system performance, based on common key metrics in the filed of RAG pipelines, assessing accuracy and reliability through various metrics and real-world scenarios. Our experimental results demonstrate significant improvements in the precision and relevance of generated responses, enhancing user experience and reducing the time required to obtain relevant answers. In summary, this paper presents a novel application of RAG pipelines and LLMs, supported by a meticulously prepared university benchmark, offering valuable insights into advanced AI techniques for academic data retrieval and setting the stage for future research in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。