arXiv:2411.03881cs.IR2024-11被引 7

用生成式大模型创造多样化查询,提升检索效果

Data Fusion of Synthetic Query Variants With Generative Large Language Models

  • 通过提示工程生成主题相关查询变体,无需标注数据
  • 在TREC新文数据集上,融合生成查询的检索效果显著优于单查询基线
  • 适合关注检索增强与大模型应用的研究者

在信息检索实验中考虑查询变异性有助于提升检索效果。基于不同主题相关查询的排序集成,通常比单一查询的排序表现更优。近期,指令微调的大语言模型在捕捉人类语言方面表现出色。本文探索了利用指令微调大模型生成合成查询变体进行数据融合的可行性。提出一种轻量级、无监督且低成本的方法,结合合理提示与数据融合技术。实验表明,当提供主题上下文时,大模型能生成更有效的查询。基于四个TREC新文基准的分析显示,基于合成查询变体的数据融合显著优于单查询基线,也优于伪相关反馈方法。代码与查询数据集已公开,供后续研究使用。

原文摘要 · Abstract (English)

Considering query variance in information retrieval (IR) experiments is beneficial for retrieval effectiveness. Especially ranking ensembles based on different topically related queries retrieve better results than rankings based on a single query alone. Recently, generative instruction-tuned Large Language Models (LLMs) improved on a variety of different tasks in capturing human language. To this end, this work explores the feasibility of using synthetic query variants generated by instruction-tuned LLMs in data fusion experiments. More specifically, we introduce a lightweight, unsupervised, and cost-efficient approach that exploits principled prompting and data fusion techniques. In our experiments, LLMs produce more effective queries when provided with additional context information on the topic. Furthermore, our analysis based on four TREC newswire benchmarks shows that data fusion based on synthetic query variants is significantly better than baselines with single queries and also outperforms pseudo-relevance feedback methods. We publicly share the code and query datasets with the community as resources for follow-up studies.

信息检索大模型查询生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。