arXiv:2506.06015cs.IR2025-06

用大模型生成新文档,提升检索效果。

On the Merits of LLM-Based Corpus Enrichment

  • 用大模型生成或修改文档,扩充语料库。
  • 新生成文档比原文档更易被检索到。
  • 适用于RAG和问答溯源任务。

生成式AI(genAI)技术——特别是大语言模型(LLMs)——与搜索之间的关系正在演变。我们提出一种新视角:利用genAI增强文档语料库,以提升基于查询的检索效果。增强方式包括修改现有文档或生成新文档。作为概念验证,我们使用LLMs生成与特定主题相关、可检索性优于现有文档的新文本。此外,我们展示了语料库增强在检索增强生成(RAG)和问答中的答案归因方面所具有的潜力。

原文摘要 · Abstract (English)

Generative AI (genAI) technologies -- specifically, large language models (LLMs) -- and search have evolving relations. We argue for a novel perspective: using genAI to enrich a document corpus so as to improve query-based retrieval effectiveness. The enrichment is based on modifying existing documents or generating new ones. As an empirical proof of concept, we use LLMs to generate documents relevant to a topic which are more retrievable than existing ones. In addition, we demonstrate the potential merits of using corpus enrichment for retrieval augmented generation (RAG) and answer attribution in question answering.

语料增强检索增强大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。