arXiv:2508.04117cs.CL2025-08被引 9

发现微调时模型过度记忆数据,影响泛化能力

Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks

  • 在微调特定阶段,模型过度记忆训练数据
  • 虽测试准确率高,但困惑度高、生成多样性差
  • 长训练、大学习率加剧问题,适合改进微调策略

预训练大语言模型通过标注数据微调以提升指令遵循能力和与人类价值观对齐。本文研究了大语言模型在推理任务微调中的学习动态,揭示了在微调特定阶段存在的未被发现的过记忆现象。在此阶段,模型过度记忆训练数据,表现出高测试困惑度但保持良好测试准确率。我们探究了导致过记忆的条件,发现该问题在多种任务、模型和微调方法中普遍存在,且长期训练和大学习率会加剧此问题。尽管过记忆模型测试准确率与正常模型相当,但其鲁棒性下降、分布外泛化能力差、生成多样性降低。基于此发现,我们提出检查点选择建议,并引入检查点融合与记忆感知重加权等技术以缓解该效应。

原文摘要 · Abstract (English)

The pretrained large language models (LLMs) are finetuned with labeled data for better instruction following ability and alignment with human values. In this paper, we study the learning dynamics of LLM finetuning on reasoning tasks and reveal the uncovered over-memorization phenomenon during a specific stage of LLM finetuning. At this stage, the LLMs have excessively memorized training data and exhibit high test perplexity while maintaining good test accuracy. We explore the conditions that contribute to over-memorization and discover that this issue is prevalent across various tasks, models, and fine-tuning methods, with prolonged training and large learning rates exacerbating the problem. Although models with over-memorization demonstrate comparable test accuracy to normal models, they suffer from reduced robustness, poor out-of-distribution generalization, and decreased generation diversity. In light of our findings on over-memorization, we offer recommendations for checkpoint selection and propose techniques such as checkpoint merging and memorization-aware reweighting to mitigate this effect.

大模型微调过记忆泛化能力生成多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。