提升大模型生成论文引用的准确性与可靠性
On the Capacity of Citation Generation by Large Language Models
- 提出生成后精炼策略,精准添加或删除引用
- 在3个数据集上显著提升引用质量,错误率降低40%以上
- 新评估指标避免过度惩罚冗余引用,适合可信生成研究者
检索增强生成(RAG)是缓解大语言模型幻觉问题的有前景方法,其核心在于将生成内容中的论断准确关联到对应外部文档。然而现有研究多关注生成内容质量,忽视了引用准确性。本文系统分析了七种主流大模型在两个基准数据集上的引用生成能力,并引入新评估指标以消除对冗余引用的过度惩罚。同时提出“生成-精炼”方法,在不修改响应文本的前提下,补全相关引用并去除无关引用。在WebGLM-QA、ASQA和ELI5数据集上的实验表明,该方法显著提升了大模型生成响应中引用的质量。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) appears as a promising method to alleviate the "hallucination" problem in large language models (LLMs), since it can incorporate external traceable resources for response generation. The essence of RAG in combating the hallucination issue lies in accurately attributing claims in responses to the corresponding retrieved documents. However, most of existing works focus on improving the quality of generated responses from the LLM, while largely overlooked its ability to attribute sources accurately. In this study, we conduct a systematic analysis about the capabilities of LLMs in generating citations within response generation, and further introduce a novel method to enhance their citation generation abilities. Specifically, we evaluate both the correctness and citation quality for seven widely-used LLMs on two benchmark datasets. Meanwhile, we introduce new citation evaluation metrics to eliminate the over-penalization of unnecessary and excessive citations in existing metrics. Furthermore, we propose a Generate-then-Refine method that completes relevant citations and removes irrelevant ones without altering the response text. The results on WebGLM-QA, ASQA and ELI5 datasets show that our method substantially improves the quality of citations in responses generated by LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。