提出新评估框架CapTrack,发现大模型微调会系统性丢失鲁棒性和默认行为
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training
- 构建能力中心化评估框架,结合行为分类与特定能力指标
- 800亿参数模型实验显示微调导致鲁棒性与默认行为显著退化
- 指令微调损伤最严重,偏好优化更保守且可部分恢复能力
大语言模型后训练能提升隐含能力、实现价值对齐、改善性能并支持领域适配。然而,后训练常引发遗忘,尤其在使用第三方预训练模型时,通常被理解为参数或事实知识的丢失。我们认为这种以准确率为重心的视角不足以刻画现代基础模型的遗忘现象,因此将遗忘定义为系统性模型漂移,导致行为和用户体验下降。为此,我们提出CapTrack——一个以能力为中心的遗忘分析框架,融合行为分类体系与面向特定能力的评估套件。基于该框架,我们在多种后训练算法、领域和模型家族(最大达800亿参数)上开展大规模实证研究。结果表明,遗忘不仅限于参数知识,还显著影响模型鲁棒性和默认行为。指令微调引发最强相对漂移,而偏好优化更为保守,并能部分恢复丢失能力。不同模型家族间差异持续存在,尚未出现通用缓解方案。
原文摘要 · Abstract (English)
Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfortunately, post-training is known to induce forgetting, especially in the ubiquitous use-case of leveraging third-party pre-trained models, which is typically understood as a loss of parametric or factual knowledge. We argue that this accuracy-centric view is insufficient for modern foundation models and instead define forgetting as systematic model drift that degrades behavior and user experience. In this context, we introduce CapTrack, a capability-centric framework for analyzing forgetting in LLMs that combines a behavioral taxonomy with an evaluation suite centered on capability-specific metrics. Using CapTrack, we conduct a large-scale empirical study across post-training algorithms, domains, and model families, including models up to 80B parameters. We find that forgetting extends beyond parametric knowledge, with pronounced drift in robustness and default behaviors. Instruction fine-tuning induces the strongest relative drift, while preference optimization is more conservative and can partially recover lost capabilities. Differences across model families persist, and no universal mitigation emerges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。