研究大模型如何高效学习罕见事实,揭示模型大小与架构对记忆稀有信息的影响。
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
- 通过标注训练数据中事实的出现频率,分析模型学习效率。
- 大模型在高频事实上表现相近,但在低频事实上差异显著。
- 适合关注模型记忆能力、训练效率的研究者参考。
样本效率是语言模型的重要属性,直接影响训练效率。现实文本中的信息呈现长尾分布,但模型需掌握高频与低频事实。高效模型更擅长在较少接触的情况下学习并保留稀有信息。本研究分析了多种不同架构和规模的模型,均在相同预训练数据上训练。通过标注关系性事实在训练语料中的频率,考察模型性能随事实频率的变化。结果表明,多数模型在高频事实上表现相似,但在低频事实上的差异明显。该分析揭示了模型架构、规模与事实学习效率之间的新关系。
原文摘要 · Abstract (English)
Sample efficiency is a crucial property of language models with practical implications for training efficiency. In real-world text, information follows a long-tailed distribution. Yet, we expect models to learn and recall frequent and infrequent facts. Sample-efficient models are better equipped to handle this challenge of learning and retaining rare information without requiring excessive exposure. This study analyzes multiple models of varying architectures and sizes, all trained on the same pre-training data. By annotating relational facts with their frequencies in the training corpus, we examine how model performance varies with fact frequency. Our findings show that most models perform similarly on high-frequency facts but differ notably on low-frequency facts. This analysis provides new insights into the relationship between model architecture, size, and factual learning efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。