用词点互信息评估检索增强生成效果,无需答案即可优化提示
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation
- 用词点互信息衡量文档与问题相关性,作为模型性能指标
- 实验显示该指标与答案准确率有显著正相关关系
- 基于该指标可优化提示选择,提升多种大模型表现
近期研究发现,大语言模型在检索增强生成任务中易受检索文档顺序影响。但目前尚无方法利用此现象提升生成质量。本文填补这一空白:提出词点互信息(PMI)可有效衡量语言模型性能,且无需预先知道问题答案。在两个问答数据集及多种大模型上的实验表明,答案准确率与词点互信息存在显著实证相关性。进一步提出两种基于文档-问题词点互信息的提示选择与构建方法,实验验证其可有效提升模型性能。
原文摘要 · Abstract (English)
Recent work suggests that large language models enhanced with retrieval-augmented generation are easily influenced by the order, in which the retrieved documents are presented to the model when solving tasks such as question answering (QA). However, there is no method to date that exploits this phenomenon to improve generation. We fill this gap. In this study, we show that the pointwise mutual information between a context and a question is an effective gauge for language model performance. Importantly, this gauge does not depend on knowing the answer to the question a priori. Through experiments on two question-answering datasets and a variety of large language models, we find evidence for an empirical correlation between answer accuracy and pointwise mutual information. Additionally, we propose two methods that use the pointwise mutual information between a document and a question as a gauge for selecting and constructing prompts that lead to better performance, whose effectiveness we demonstrate through experimentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。