大模型事实记忆能力随模型规模和话题频率提升,且可被定量预测。
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency

- 用模型参数量与训练数据中话题频率的对数线性组合建模记忆表现。
- 该模型解释了16个密集模型60%的差异,家族内可达74%-94%。
- 适用于评估模型事实记忆能力,尤其关注知识密度的研究者。
尽管缩放定律能描述大型语言模型的整体性能,但尚未有定律将事实记忆能力同时与模型规模和训练数据构成关联。我们对38个模型在超过8,900个学术引用上进行了评估,使用自动化引用验证系统。记忆质量在模型参数量与训练数据中话题表征的对数线性组合下呈现饱和型曲线(sigmoid)。这两个变量单独解释了来自四个模型家族的16个密集模型中60%的方差,而在单一家族内部可达到74%-94%。该形式符合一种基于超叠加的解释:回忆能力受信号-噪声比调控,信号强度随概念频率增加,噪声基线随模型容量上升。
原文摘要 · Abstract (English)
While scaling laws govern aggregate large language model performance, no scaling law has linked factual recall to both model size and training-data composition. We evaluated 38 models on over 8,900 scholarly references evaluated by an automated reference verification system. Recall quality follows a sigmoid in the log-linear combination of model parameter count and topic representation in training data. These two variables alone explain 60% of the variance across 16 dense models from four families, rising to 74-94% within individual families. The form matches a superposition-inspired account in which recall is gated by a signal-to-noise ratio: signal strength scales with concept frequency and the noise floor with model capacity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。