arXiv:2502.20186cs.CLcs.LG2025-02EMNLP被引 5

通过分层加权任务向量,分离特定任务与通用指令理解知识。

Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge

  • 按层分配权重,区分任务特异性与指令遵循成分。
  • 在多个基准上提升多任务学习与选择性遗忘效果。
  • 适合需要高效模型合并与编辑的场景。

大语言模型(LLMs)通过微调展现强任务特定能力,但合并多个微调模型常因指令遵循组件重叠导致性能下降。任务算术(TA)通过组合微调生成的任务向量实现多任务学习与任务遗忘,但难以分离任务特异性知识与通用指令遵循行为。为此,我们提出分层感知任务算术(LATA),根据任务向量与指令遵循或任务特异性组件的对齐程度,为各层分配特定权重。通过增强任务相关层、抑制指令遵循层,LATA在保持整体模型性能的同时,显著提升任务学习与遗忘表现。在WikiText-2、GSM8K和HumanEval等多个基准上的实验表明,LATA优于现有方法,实现更高任务准确率与更优对齐度,且输出质量下降最小。研究揭示了分层分析在解耦任务特异与通用知识中的关键作用,为高效模型合并与编辑提供了稳健框架。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to overlapping instruction-following components. Task Arithmetic (TA), which combines task vectors derived from fine-tuning, enables multi-task learning and task forgetting but struggles to isolate task-specific knowledge from general instruction-following behavior. To address this, we propose Layer-Aware Task Arithmetic (LATA), a novel approach that assigns layer-specific weights to task vectors based on their alignment with instruction-following or task-specific components. By amplifying task-relevant layers and attenuating instruction-following layers, LATA improves task learning and forgetting performance while preserving overall model utility. Experiments on multiple benchmarks, including WikiText-2, GSM8K, and HumanEval, demonstrate that LATA outperforms existing methods in both multi-task learning and selective task forgetting, achieving higher task accuracy and alignment with minimal degradation in output quality. Our findings highlight the importance of layer-wise analysis in disentangling task-specific and general-purpose knowledge, offering a robust framework for efficient model merging and editing.

模型合并任务分离LLM编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。