arXiv:2504.01241cs.CL2025-04被引 9

对比小模型在持续学习中遗忘程度,找出现有最佳适应者。

Catastrophic Forgetting in LLMs: A Comparative Analysis Across Language Tasks

  • 用提示工程和任务调整测试多模型持续微调能力。
  • Phi-3.5-mini 忘记最少且学习能力强,表现最优。
  • 适合需要长期学习新任务的智能体系统参考。

大型语言模型(LLMs)在自然语言理解(NLU)任务中取得显著进展。随着基于LLM的智能体走向自主处理专业化任务的未来,模型在学习新任务时避免遗忘旧知识的能力——即灾难性遗忘——变得至关重要。本研究评估了多个参数量小于100亿的开源LLM在GLUE基准中的关键NLU任务(包括SST-2、MRPC、CoLA和MNLI)上的持续微调表现。通过提示工程与任务特定调整,比较各模型在保持已有知识的同时学习新任务的能力。结果表明,Phi-3.5-mini表现出极低的遗忘率并具备强大学习能力,非常适合持续学习场景。Orca-2-7b与Qwen2.5-7B则在微调后展现出优异的学习能力和整体性能。本研究深化了对LLMs中灾难性遗忘的理解,并强调提示工程在优化持续学习性能中的作用。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have significantly advanced Natural Language Processing (NLP), particularly in Natural Language Understanding (NLU) tasks. As we progress toward an agentic world where LLM-based agents autonomously handle specialized tasks, it becomes crucial for these models to adapt to new tasks without forgetting previously learned information - a challenge known as catastrophic forgetting. This study evaluates the continual fine-tuning of various open-source LLMs with different parameter sizes (specifically models under 10 billion parameters) on key NLU tasks from the GLUE benchmark, including SST-2, MRPC, CoLA, and MNLI. By employing prompt engineering and task-specific adjustments, we assess and compare the models' abilities to retain prior knowledge while learning new tasks. Our results indicate that models such as Phi-3.5-mini exhibit minimal forgetting while maintaining strong learning capabilities, making them well-suited for continual learning environments. Additionally, models like Orca-2-7b and Qwen2.5-7B demonstrate impressive learning abilities and overall performance after fine-tuning. This work contributes to understanding catastrophic forgetting in LLMs and highlights prompting engineering to optimize model performance for continual learning scenarios.

大模型持续学习遗忘问题提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。