发现大模型查询扩展可能因数据泄露而虚高表现
Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
- 用事实验证任务检验大模型生成文档是否含真实证据信息
- 生成内容若包含真证据句子,性能提升显著平均达12.3%
- 提醒研究者警惕基准数据泄露对评估结果的干扰
基于大语言模型(LLMs)的查询扩展方法在零样本检索任务中表现出色。这些方法假设大模型能生成假想文档,融入查询向量后可提升真实证据的检索效果。然而,我们质疑这一假设,通过事实验证任务分析生成文档是否包含由真实证据蕴含的信息,并评估其对性能的影响。结果表明,当生成文档包含与真实证据蕴含的句子时,性能提升始终显著。这暗示事实验证基准中可能存在知识泄露,导致大模型查询扩展方法的表现被高估。
原文摘要 · Abstract (English)
Query expansion methods powered by large language models (LLMs) have demonstrated effectiveness in zero-shot retrieval tasks. These methods assume that LLMs can generate hypothetical documents that, when incorporated into a query vector, enhance the retrieval of real evidence. However, we challenge this assumption by investigating whether knowledge leakage in benchmarks contributes to the observed performance gains. Using fact verification as a testbed, we analyze whether the generated documents contain information entailed by ground-truth evidence and assess their impact on performance. Our findings indicate that, on average, performance improvements consistently occurred for claims whose generated documents included sentences entailed by gold evidence. This suggests that knowledge leakage may be present in fact-verification benchmarks, potentially inflating the perceived performance of LLM-based query expansion methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。