FIM预训练让模型更易记住片段,且记忆强度随重复次数线性增长。
Memorization Dynamics of Fill-in-the-Middle Pretraining

- 用前缀探针研究FIM训练下模型的记忆行为
- 记忆程度随重复次数近似线性上升,短片段更易复现
- 仅测单一跨度或格式会遗漏关键记忆特性
填空中间(FIM)是一种广泛用于赋予因果语言模型补全能力的预训练目标,但其对原文记忆的影响尚未充分探索。我们通过在包含重复古腾堡文本的FineWeb-Gutenberg语料上,对匹配的Llama 3.2模型分别采用FIM和标准左到右(LTR)目标进行预训练,控制条件下研究FIM的记忆动态。通过前缀探针发现,FIM更倾向于恢复短或部分匹配的片段,而LTR更常对长且精确的延续给出高置信度。在测试范围内,FIM训练下的原文提取率随重复次数近似线性增长。评估原生FIM格式探针表明,仅靠后缀上下文不足以支持记忆:FIM训练下的原文召回仍强烈依赖前缀上下文。结果还显示,仅评估单一跨度长度或探针格式可能遗漏记忆行为的重要细节。
原文摘要 · Abstract (English)
Fill-in-the-middle (FIM) is a pretraining objective widely used to equip causal language models with infilling ability, yet its effect on verbatim memorization remains underexplored. We study the memorization dynamics of FIM in a controlled setting by pretraining matched Llama 3.2 models with FIM and standard left-to-right (LTR) objectives on a FineWeb-Gutenberg corpus containing repeated Gutenberg excerpts. With prefix-based probes, FIM more often recovers short or partially matching spans, while LTR more often assigns high confidence to long exact continuations. We observe that verbatim extraction under FIM-training grows approximately linearly with repetitions over the tested range. Evaluating native FIM-format probes reveals that suffix context is not sufficient: verbatim recall under FIM-training remains strongly anchored in prefix context. Our results also show that evaluating only one span length or probing format can miss important nuances in memorization behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。