发现大模型预训练中存在遗忘现象,并提出更优检测方法。
Exploring Forgetting in Large Language Model Pre-Training
- 用新指标替代困惑度,更好检测实体记忆保留情况。
- 验证了预训练阶段确实存在遗忘,且学习曲线揭示动态规律。
- 提出低成本缓解方案,适合关注模型长期记忆的研究者。
灾难性遗忘仍是构建全能大语言模型的主要障碍。尽管已有研究关注微调阶段的任务级遗忘,但对预训练过程中的遗忘关注甚少。本文系统探索了预训练中遗忘的存在性与测量方式,质疑传统困惑度(PPL)等指标,引入新指标以更准确检测实体记忆保留。基于修正后的评估框架,我们提出了低成本、简单的预训练阶段遗忘缓解方法。进一步通过细致分析学习曲线,揭示了遗忘的动态机制。大规模评估与分析为未来大模型研究提供了重要参考。
原文摘要 · Abstract (English)
Catastrophic forgetting remains a formidable obstacle to building an omniscient model in large language models (LLMs). Despite the pioneering research on task-level forgetting in LLM fine-tuning, there is scant focus on forgetting during pre-training. We systematically explored the existence and measurement of forgetting in pre-training, questioning traditional metrics such as perplexity (PPL) and introducing new metrics to better detect entity memory retention. Based on our revised assessment of forgetting metrics, we explored low-cost, straightforward methods to mitigate forgetting during the pre-training phase. Further, we carefully analyzed the learning curves, offering insights into the dynamics of forgetting. Extensive evaluations and analyses on forgetting of pre-training could facilitate future research on LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。