大模型检索能力随预训练算力增长而提升,且与上下文学习表现强相关。
Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- 通过对比不同规模模型的检索表现,发现性能随参数量、训练时长和预训练算力同步提升。
- 零样本BEIR任务中,70亿参数模型在2万亿令牌数据上训练后检索准确率显著优于小模型。
- 模型上下文学习能力与检索能力高度相关,适合构建基于LLM的检索系统开发者参考。
检索性能如何随预训练算力变化?我们评估了从1.25亿到70亿参数的大型语言模型在10亿到超过2万亿令牌数据上的检索表现。结果表明,零样本BEIR任务中的检索性能可预测地随模型规模、训练时长和估算的浮点运算量(FLOPs)增长。此外,在各类检索任务中,上下文学习得分与检索得分呈强相关性。研究揭示了构建基于大语言模型的检索器的关键启示。
原文摘要 · Abstract (English)
How does retrieval performance scale with pretraining FLOPs? We benchmark retrieval performance across LLM model sizes from 125 million parameters to 7 billion parameters pretrained on datasets ranging from 1 billion tokens to more than 2 trillion tokens. We find that retrieval performance on zero-shot BEIR tasks predictably scales with LLM size, training duration, and estimated FLOPs. We also show that In-Context Learning scores are strongly correlated with retrieval scores across retrieval tasks. Finally, we highlight the implications this has for the development of LLM-based retrievers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。