用合成数据提升领域专用生成式检索模型的训练效率与效果
On Synthetic Data Strategies for Domain-Specific Generative Retrieval
- 通过大模型生成多粒度查询,融合领域搜索约束
- 基于初始模型预测挖掘难负例,优化文档排序
- 在多个领域公开数据集上验证了方法有效性
本文研究在构建领域专用语料库的生成式检索模型时,合成数据生成策略的有效性,以应对人工标注领域内查询的可扩展性挑战。针对两阶段训练框架:第一阶段聚焦于从查询中解码文档标识符,探索了大模型生成的多粒度(如段落、句子)查询及领域相关的搜索约束,以更好地捕捉细微的相关性信号;第二阶段旨在通过偏好学习优化文档排序,研究基于初始模型预测挖掘难负例的策略。在多个领域的公开数据集上的实验表明,所提出的合成数据生成与难负例采样方法具有显著有效性。
原文摘要 · Abstract (English)
This paper investigates synthetic data generation strategies in developing generative retrieval models for domain-specific corpora, thereby addressing the scalability challenges inherent in manually annotating in-domain queries. We study the data strategies for a two-stage training framework: in the first stage, which focuses on learning to decode document identifiers from queries, we investigate LLM-generated queries across multiple granularity (e.g. chunks, sentences) and domain-relevant search constraints that can better capture nuanced relevancy signals. In the second stage, which aims to refine document ranking through preference learning, we explore the strategies for mining hard negatives based on the initial model's predictions. Experiments on public datasets over diverse domains demonstrate the effectiveness of our synthetic data generation and hard negative sampling approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。