arXiv:2410.20008cs.CLcs.LG2024-10EMNLP被引 17

揭秘指令微调后大模型如何分层存储多任务知识。

Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models

  • 用矩阵分析法对比预训练与指令微调模型的表征差异。
  • 发现60多个任务中部分需微调才能有效编码知识。
  • 定位到模型从通用到任务专属表征的转变层,指导高效迁移学习。

在超过60个自然语言处理任务上,研究预训练大语言模型(LLMs)中任务特定信息的编码方式及指令微调对其表征的影响。通过矩阵分析工具,对比预训练与指令微调模型的表征差异。结果表明,部分任务在预训练阶段已具备编码能力,而多数任务则显著受益于指令微调。进一步识别出模型在特定层级由高层通用表示向任务导向表示过渡。该发现深化了对大模型工作机制的理解,为参数高效迁移学习和多任务学习提供了新思路。

原文摘要 · Abstract (English)

Fine-tuning pre-trained large language models (LLMs) on a diverse array of tasks has become a common approach for building models that can solve various natural language processing (NLP) tasks. However, where and to what extent these models retain task-specific knowledge remains largely unexplored. This study investigates the task-specific information encoded in pre-trained LLMs and the effects of instruction tuning on their representations across a diverse set of over 60 NLP tasks. We use a set of matrix analysis tools to examine the differences between the way pre-trained and instruction-tuned LLMs store task-specific information. Our findings reveal that while some tasks are already encoded within the pre-trained LLMs, others greatly benefit from instruction tuning. Additionally, we pinpointed the layers in which the model transitions from high-level general representations to more task-oriented representations. This finding extends our understanding of the governing mechanisms of LLMs and facilitates future research in the fields of parameter-efficient transfer learning and multi-task learning.

大模型多任务学习表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。