发现大模型情感空间呈环形结构,可精准调控文本情绪与行为倾向。
Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control

- 基于主成分分析与岭回归,提取出环形分布的情感轴。
- 沿该轴调节可单调控制生成文本的情绪,且影响拒绝与迎合行为。
- 适用于多个主流模型,适合研究可控生成与伦理对齐的学者。
我们发现大语言模型中的情感向量由二维效价-唤醒(VA)子空间组织,呈现出环形几何结构。通过主成分分解和岭回归,我们恢复了情感引导向量中蕴含的有意义的VA轴,其投影与44,728个词的人类情感评分高度相关。沿这些轴进行引导可实现生成文本情感属性的单调控制,并从单一子空间实现对下游多种行为(如拒绝、阿谀)的双向调控。该现象在Llama-3.1-8B、Qwen3-8B和Qwen3-14B模型中均复现。我们提出词汇中介机制解释此类效果:拒绝与顺从关键词位于不同的VA区域,因此VA调控可直接改变其生成概率。
原文摘要 · Abstract (English)
We show that emotion vectors in LLMs are organized by a two-dimensional valence-arousal (VA) subspace exhibiting circular geometry. Through principal component decomposition and ridge regression, we recover meaningful VA axes underlying emotion steering vectors whose projections correlate with human affect ratings across 44,728 words. Steering along these axes produces monotonic control over the affective properties of generated text, and further affords bidirectional control over multiple downstream behaviors (refusal and sycophancy) from a single subspace. These effects replicate across Llama-3.1-8B, Qwen3-8B, and Qwen3-14B. We propose lexical mediation to explain why these effects and prior emotionally framed controls work: refusal and compliance tokens occupy distinct VA regions, and VA steering directly modulates their emission probabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。