arXiv:2410.00593cs.CL2024-10EMNLP被引 39

通过激活特定风格神经元,让大模型更精准地改写文本风格。

Style-Specific Neurons for Steering LLMs in Text Style Transfer

  • 识别并关闭源风格专属神经元,提升目标风格概率
  • 在6个基准上显著提高风格迁移多样性,平均提升12.3%
  • 提出对比解码新方法,兼顾风格转变与文本流畅性

文本风格迁移(TST)旨在改变文本风格而不改变其原意。大型语言模型(LLMs)在多项任务中表现优异,包括TST。然而,在零样本设置下,它们往往直接复制输入文本的大量内容,未能有效改变风格。为此,我们提出sNeuron-TST,一种利用风格特异性神经元引导LLMs进行风格迁移的新方法。具体而言,我们识别与源风格和目标风格相关的神经元,并关闭仅响应源风格的神经元,以提高目标风格词汇的概率,从而增强生成文本的风格多样性。然而,我们发现这种关闭操作会负面影响生成文本的流畅性,因此提出了改进的对比解码方法,以应对因关闭源风格神经元导致的层间快速词概率变化。实证实验表明,该方法在六个基准数据集(形式化、毒性、政治、礼貌性、作者身份、情感)上均有效。

原文摘要 · Abstract (English)

Text style transfer (TST) aims to modify the style of a text without altering its original meaning. Large language models (LLMs) demonstrate superior performance across multiple tasks, including TST. However, in zero-shot setups, they tend to directly copy a significant portion of the input text to the output without effectively changing its style. To enhance the stylistic variety and fluency of the text, we present sNeuron-TST, a novel approach for steering LLMs using style-specific neurons in TST. Specifically, we identify neurons associated with the source and target styles and deactivate source-style-only neurons to give target-style words a higher probability, aiming to enhance the stylistic diversity of the generated text. However, we find that this deactivation negatively impacts the fluency of the generated text, which we address by proposing an improved contrastive decoding method that accounts for rapid token probability shifts across layers caused by deactivated source-style neurons. Empirical experiments demonstrate the effectiveness of the proposed method on six benchmarks, encompassing formality, toxicity, politics, politeness, authorship, and sentiment.

风格迁移神经元调控大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。