arXiv:2411.03675cs.CLcs.AI2024-11被引 1

提升大模型引用生成能力,解决虚构引用和质量不足问题。

QUILL: Quotation Generation Enhancement of Large Language Models

  • 构建包含32022条引用的双语知识库,覆盖多维度内容。
  • 设计专属重排指标,使检索结果更贴近人类偏好。
  • 评估体系涵盖五项标准,与人工评分高度一致。

尽管大型语言模型已成为优秀的写作助手,但在引用生成方面仍存在困难:要么编造事实性引文,要么无法产出超越人类预期的引述。为弥合这一差距,本文系统研究了引用生成任务的评估与改进方法。首先建立了一个全面且自动化的评估体系,包含五个维度及对应的自动评价指标。为提升模型表现,构建了一个涵盖32,022条引文的广度与深度兼具的双语知识库;并基于上述标准设计专用重排指标,对知识库中检索到的引文进行排序优化。大量实验表明,所提指标与人类偏好高度相关。现有大模型在生成理想引文方面表现不佳,而本研究的知识库与重排机制有效缩小了这一差距。数据集与代码已开源于https://github.com/GraceXiaoo/QUILL。

原文摘要 · Abstract (English)

While Large language models (LLMs) have become excellent writing assistants, they still struggle with quotation generation. This is because they either hallucinate when providing factual quotations or fail to provide quotes that exceed human expectations. To bridge the gap, we systematically study how to evaluate and improve LLMs' performance in quotation generation tasks. We first establish a holistic and automatic evaluation system for quotation generation task, which consists of five criteria each with corresponding automatic metric. To improve the LLMs' quotation generation abilities, we construct a bilingual knowledge base that is broad in scope and rich in dimensions, containing up to 32,022 quotes. Moreover, guided by our critiria, we further design a quotation-specific metric to rerank the retrieved quotations from the knowledge base. Extensive experiments show that our metrics strongly correlate with human preferences. Existing LLMs struggle to generate desired quotes, but our quotation knowledge base and reranking metric help narrow this gap. Our dataset and code are publicly available at https://github.com/GraceXiaoo/QUILL.

引用生成知识库大模型评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。