arXiv:2603.13249cs.CLcs.AI2026-03被引 3

定位三个关键注意力头,实现精准人格控制且不破坏文本连贯性

Steering at the Source: Style Modulation Heads for Robust Persona Control

  • 通过几何分析定位仅3个控制人格与风格的注意力头
  • 仅干预特定头时,人格控制更稳定,连贯性下降减少60%以上
  • 适合需要安全可控生成的对话系统开发者

激活引导为无需微调即可控制大语言模型提供了高效机制。尽管能有效控制目标特征(如人格),但连贯性退化仍是影响安全性和实际部署的主要障碍。我们假设该问题源于对残差流的干预,其无差别影响聚合特征并放大非目标噪声。本文识别出仅三个注意力头构成的稀疏子集,独立控制人格与风格形成,称为风格调制头。这些头可通过层内余弦相似度与头级贡献度结合的几何分析定位。实验证明,仅针对这些特定头进行干预,可在显著缓解残差流引导中出现的连贯性退化的同时,实现稳健的行为控制。更广泛而言,精确的组件级定位使模型控制更安全、更精准。

原文摘要 · Abstract (English)

Activation steering offers a computationally efficient mechanism for controlling Large Language Models (LLMs) without fine-tuning. While effectively controlling target traits (e.g., persona), coherency degradation remains a major obstacle to safety and practical deployment. We hypothesize that this degradation stems from intervening on the residual stream, which indiscriminately affects aggregated features and inadvertently amplifies off-target noise. In this work, we identify a sparse subset of attention heads (only three heads) that independently govern persona and style formation, which we term Style Modulation Heads. Specifically, these heads can be localized via geometric analysis of internal representations, combining layer-wise cosine similarity and head-wise contribution scores. We demonstrate that intervention targeting only these specific heads achieves robust behavioral control while significantly mitigating the coherency degradation observed in residual stream steering. More broadly, our findings show that precise, component-level localization enables safer and more precise model control.

模型控制注意力头人格生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。