arXiv:2505.14309cs.CL2025-05EMNLP被引 1

通过增加查询与检索内容重叠,可提升语言模型训练效率40%。

Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency

  • 用改写查询生成合成上下文,主动提高查询与检索内容重叠度。
  • 超过临界阈值后,测试困惑度下降,学习速度加快40%。
  • 适合关注模型训练加速与数据效率的研究者。

检索增强型语言模型在计算资源更少的情况下,性能可媲美更大模型。其效果关键取决于查询与检索上下文的重叠程度,但最优重叠度尚未明确。本文系统研究不同查询-上下文重叠水平对模型训练和推理的影响。实验表明,初始阶段重叠度提升影响较小,但超过临界阈值后,测试困惑度显著降低,模型学习速度明显加快。基于此发现,我们证明通过改写查询生成合成上下文,可使数据效率提升,训练时间减少约40%,且不损失性能。我们在问答任务上验证了困惑度结果的有效性,确认检索增强方法在实际应用中仍具优势。研究为语言模型预训练中检索机制的优化提供了实证依据。

原文摘要 · Abstract (English)

Retrieval-augmented language models have demonstrated performance comparable to much larger models while requiring fewer computational resources. The effectiveness of these models crucially depends on the overlap between query and retrieved context, but the optimal degree of this overlap remains unexplored. In this paper, we systematically investigate how varying levels of query--context overlap affect model performance during both training and inference. Our experiments reveal that increased overlap initially has minimal effect, but substantially improves test-time perplexity and accelerates model learning above a critical threshold. Building on these findings, we demonstrate that deliberately increasing overlap through synthetic context can enhance data efficiency and reduce training time by approximately 40\% without compromising performance. We specifically generate synthetic context through paraphrasing queries. We validate our perplexity-based findings on question-answering tasks, confirming that the benefits of retrieval-augmented language modeling extend to practical applications. Our results provide empirical evidence of significant optimization potential for retrieval mechanisms in language model pretraining.

检索增强训练效率语言模型数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。