通过激活工程让大模型更像人一样表达情绪。
Conversations: Love Them, Hate Them, Steer Them
- 用归因修补定位关键激活组件,找到情绪调控切入点。
- 对比正负情感文本生成情绪向量,提升回应的喜悦与信任感。
- 无需大量微调,适合想改进对话情感表达的研究者。
大型语言模型(LLMs)的对话能力日益流畅,但赋予其细腻、类人的情感表达仍是重大挑战。现有对齐方法多关注表面输出或需大量微调。本文表明,针对性激活工程可使LLaMA 3.1-8B展现更类人的感情细节。我们首先使用归因修补识别因果关键组件,在诊断性对话任务中观察激活模式,定位干预点。随后,从对比文本对(正/负例)的激活差异中提取情感表达向量。将这些向量应用于新对话提示后,响应显著增强情感特征:正向情感(如喜悦、信任)上升,第一人称代词使用频率提高,体现更强个人参与感。研究提供一种精确且可解释的情绪属性控制方法,有助于构建更对齐、更具同理心的对话AI。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate increasing conversational fluency, yet instilling them with nuanced, human-like emotional expression remains a significant challenge. Current alignment techniques often address surface-level output or require extensive fine-tuning. This paper demonstrates that targeted activation engineering can steer LLaMA 3.1-8B to exhibit more human-like emotional nuances. We first employ attribution patching to identify causally influential components, to find a key intervention locus by observing activation patterns during diagnostic conversational tasks. We then derive emotional expression vectors from the difference in the activations generated by contrastive text pairs (positive vs. negative examples of target emotions). Applying these vectors to new conversational prompts significantly enhances emotional characteristics: steered responses show increased positive sentiment (e.g., joy, trust) and more frequent first-person pronoun usage, indicative of greater personal engagement. Our findings offer a precise and interpretable method for controlling specific emotional attributes in LLMs, contributing to developing more aligned and empathetic conversational AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。