arXiv:2502.06115cs.CLcs.LG2025-02NAACL被引 3

提出分层激活干预方法,提升大模型任务适应效率。

Task-driven Layerwise Additive Activation Intervention

  • 按层添加可学习激活项,动态优化生成过程。
  • 在多个数据集上显著提升预训练模型准确率。
  • 无需大量提示词,样本效率优于现有方法。

现代语言模型在自然语言处理生成任务中取得显著进展,但在实时应用中适应新场景仍存挑战。激活干预通过识别并操纵模型中间激活来引导生成,但现有方法依赖启发式规则或需大量提示才能确定有效干预。本文提出一种分层加性激活干预框架,优化干预流程,提升样本效率。我们在多个数据集上进行了基准测试,结果表明该框架能有效提升预训练模型的准确性,并优于现有干预基线。

原文摘要 · Abstract (English)

Modern language models (LMs) have significantly advanced generative modeling in natural language processing (NLP). Despite their success, LMs often struggle with adaptation to new contexts in real-time applications. A promising approach to task adaptation is activation intervention, which steers the LMs' generation process by identifying and manipulating the activations. However, existing interventions are highly dependent on heuristic rules or require many prompt inputs to determine effective interventions. This paper proposes a layer-wise additive activation intervention framework that optimizes the intervention process, thus enhancing the sample efficiency. We benchmark our framework on various datasets, demonstrating improvements in the accuracy of pre-trained LMs and competing intervention baselines.

语言模型激活干预高效适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。