arXiv:2604.04064cs.CLcs.AI2026-04被引 2

首次对比小模型情感表征提取方法,发现生成式更优且可精准控制。

Extracting and Steering Emotion Representations in Small Language Models: A Methodological Comparison

  • 用生成与理解两类方法提取小模型情感向量,验证其有效性。
  • 情感表征位于中间层(约50%深度),呈通用U型分布。
  • 可实现精准情感操控,但多语言下存在安全风险。

参数量在100M-10B范围的小型语言模型日益用于生产系统,但其是否具备前沿模型中发现的内部情感表征尚不明确。本文首次对小模型的情感向量提取方法进行对比分析,评估了9个模型、5种架构(GPT-2、Gemma、Qwen、Llama、Mistral),涵盖20种情绪,使用两种提取方式(基于生成和基于理解)。基于生成的提取方法在情感分离上表现更优(Mann-Whitney p = 0.007;Cohen's d = -107.5),优势受指令微调和架构影响。情感表示集中在中间变压器层(约50%深度),呈现架构无关的U形曲线,适用于124M至3B参数模型。通过4个模型的表示各向异性基线验证,并经外部情感分类器独立验证(92%成功率,37/40场景),确认因果行为效应。操控实验揭示三种模式——精准转换(外科式)、重复坍缩与爆炸性退化,由模型架构而非规模决定。在Qwen中观察到跨语言情感纠缠,操控激活语义一致的中文词元,而强化学习人类反馈(RLHF)无法抑制,提示多语言部署中的安全风险。本研究为开源权重模型的情感研究提供方法指南,推动模型医学系列进展,连接外部行为分析与内部表征研究。

原文摘要 · Abstract (English)

Small language models (SLMs) in the 100M-10B parameter range increasingly power production systems, yet whether they possess the internal emotion representations recently discovered in frontier models remains unknown. We present the first comparative analysis of emotion vector extraction methods for SLMs, evaluating 9 models across 5 architectural families (GPT-2, Gemma, Qwen, Llama, Mistral) using 20 emotions and two extraction methods (generation-based and comprehension-based). Generation-based extraction produces statistically superior emotion separation (Mann-Whitney p = 0.007; Cohen's d = -107.5), with the advantage modulated by instruction tuning and architecture. Emotion representations localize at middle transformer layers (~50% depth), following a U-shaped curve that is architecture-invariant from 124M to 3B parameters. We validate these findings against representational anisotropy baselines across 4 models and confirm causal behavioral effects through steering experiments, independently verified by an external emotion classifier (92% success rate, 37/40 scenarios). Steering reveals three regimes -- surgical (coherent text transformation), repetitive collapse, and explosive (text degradation) -- quantified by perplexity ratios and separated by model architecture rather than scale. We document cross-lingual emotion entanglement in Qwen, where steering activates semantically aligned Chinese tokens that RLHF does not suppress, raising safety concerns for multilingual deployment. This work provides methodological guidelines for emotion research on open-weight models and contributes to the Model Medicine series by bridging external behavioral profiling with internal representational analysis.

情感表征小模型可控生成多语言安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。