arXiv:2502.11181cs.IRcs.AI2025-02被引 12

用概念覆盖度指导生成查询,让科研文档检索更全面。

Improving Scientific Document Retrieval with Concept Coverage-based Query Set Generation

  • 根据已生成查询的遗漏概念动态调整后续生成
  • 在多个数据集上显著提升检索准确率
  • 适合需要高覆盖率的科研文献检索场景

在科学领域,构建大规模人工标注数据集面临巨大挑战,因需专业知识。近期方法使用大语言模型生成合成查询,作为真实用户查询的代理,但难以控制生成内容,常导致文档中学术概念覆盖不全。本文提出概念覆盖度驱动的查询集生成框架CCQGen,旨在生成涵盖文档全部概念的查询集。其核心在于:根据已有查询的覆盖情况,识别未充分覆盖的概念,并将其作为后续查询生成的条件,引导新查询补充前序查询的不足,实现对文档的全面理解。大量实验表明,CCQGen显著提升了查询质量和检索性能。

原文摘要 · Abstract (English)

In specialized fields like the scientific domain, constructing large-scale human-annotated datasets poses a significant challenge due to the need for domain expertise. Recent methods have employed large language models to generate synthetic queries, which serve as proxies for actual user queries. However, they lack control over the content generated, often resulting in incomplete coverage of academic concepts in documents. We introduce Concept Coverage-based Query set Generation (CCQGen) framework, designed to generate a set of queries with comprehensive coverage of the document's concepts. A key distinction of CCQGen is that it adaptively adjusts the generation process based on the previously generated queries. We identify concepts not sufficiently covered by previous queries, and leverage them as conditions for subsequent query generation. This approach guides each new query to complement the previous ones, aiding in a thorough understanding of the document. Extensive experiments demonstrate that CCQGen significantly enhances query quality and retrieval performance.

文档检索查询生成LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。