arXiv:2409.15820cs.LGcs.CL2024-09被引 6

发现大模型微调时通过激活特定注意力头快速适配任务

Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns

  • 通过梯度分析揭示微调中注意力头的选择性激活机制
  • 复杂任务的激活模式由基础任务模式组合而成
  • 少量样本微调即可显著改变注意力模式,提升效率

大语言模型在复杂任务上的表现仍不理想。主要原因在于当前模型依赖数据驱动学习,而复杂任务的指令稀缺且难以构建。相反,模型在简单任务上可快速学习,因预训练阶段已积累充足先验知识。若能揭示这种快速泛化的前提与机制,将显著提升模型学习复杂任务的效率与效果。本文采用基于梯度的方法,从注意力模式视角剖析监督微调(SFT)如何使模型适配下游任务。研究发现:(1) 模型在SFT过程中选择性激活特定任务的注意力头;(2) 复杂任务的激活模式是基础任务模式的组合;(3) 少量参数调整即可在少量样本上显著改变激活模式。基于此,实验验证了该机制对SFT效率与效果的提升作用。

原文摘要 · Abstract (English)

LLMs' performance on complex tasks is still unsatisfactory. A key issue is that presently LLMs learn in a data-driven schema, while the instructions about these complex tasks are both scarce and hard to collect or construct. On the contrary, a prominent phenomenon is that LLMs can learn rather fast on simpler tasks with adequate prior knowledge captured during pretraining stage. Thus, if the prerequisite and mechanism of such rapid generalization could be elucidated, it could enhance the efficiency and effectiveness of the LLM's ability to learn complex tasks. Thus, in this paper, we employ a gradient-based method, to dissect the process that the SFT process adapts LLMs to downstream tasks via the perspective of attention patterns. We find that: (1) LLMs selectively activate task-specific attention heads during SFT; (2) activation patterns for complex tasks are combinations of basic task patterns; and (3) changes in a few parameters can significantly impact activation patterns after SFT on a small number of samples.Based on these insights, experiments are conducted to actually enhance the efficiency and effectiveness of SFT.

大模型微调注意力机制任务适配高效学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。