通过激活修补解析大模型如何基于人格进行推理
Dissecting Persona-Driven Reasoning in Language Models via Activation Patching

- 用激活修补技术分析模型中人格信息的编码机制
- 早期MLP层将人格标记转化为丰富语义表示
- 特定注意力头过度关注种族与肤色等身份特征
大型语言模型(LLMs)展现出在不同人格间灵活切换的能力。本研究探讨赋予人格如何影响模型在客观任务上的推理表现。通过激活修补技术,我们首次揭示模型中关键组件如何编码人格相关信息。研究发现,早期多层感知机(MLP)层不仅关注输入的句法结构,还处理其语义内容,将人格标记转化为更丰富的表征,进而由中间多头注意力(MHA)层用于塑造输出。此外,我们识别出若干注意力头显著关注种族和颜色相关身份特征。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit remarkable versatility in adopting diverse personas. In this study, we examine how assigning a persona influences a model's reasoning on an objective task. Using activation patching, we take a first step toward understanding how key components of the model encode persona-specific information. Our findings reveal that the early Multi-Layer Perceptron (MLP) layers attend not only to the syntactic structure of the input but also process its semantic content. These layers transform persona tokens into richer representations, which are then used by the middle Multi-Head Attention (MHA) layers to shape the model's output. Additionally, we identify specific attention heads that disproportionately attend to racial and color-based identities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。