小模型情绪表征具高度一致性,行为差异源于共享表征之上的层次。
Shared Emotion Geometry Across Small Language Models: A Cross-Architecture Study of Representation, Behavior, and Methodological Confounds

- 统一流程提取12个模型的21类情绪向量,分析其几何结构相似性
- 5个成熟模型情绪几何相似度极高(斯皮尔曼相关0.74-0.92)
- 揭示方法混淆效应,强调需分层解析实验差异
我们在统一理解模式流程下,以fp16精度从12个小型语言模型(六种架构×base/instruct,参数量1B-8B)中提取21类情绪向量集,并通过原始余弦RDM的表示相似性分析比较其几何结构。五个成熟架构(Qwen 2.5 1.5B、SmolLM2 1.7B、Llama 3.2 3B、Mistral 7B v0.3、Llama 3.1 8B)共享几乎相同的21类情绪几何结构,成对RDM斯皮尔曼相关系数为0.74–0.92。该普适性在行为特征截然相反的模型间依然成立:尽管Qwen 2.5与Llama 3.2在MTI合规性上处于对立极点,其情绪RDM相关性仍达rho = 0.81,表明行为差异出现在共享情绪表征之上。唯一未成熟的模型Gemma-3 1B base表现出极端残差流各向异性(0.997),且所有几何描述符均受RLHF重构;而其余五组成熟家族内部base与instruct版本的RDM相关性均≥0.92(Mistral 7B v0.3达0.985),表明仅尚未组织化的表征会被RLHF重构。方法上,我们发现先前研究误读的“理解-生成”方法效应实则分解为四层:粗粒度方法差异、生成阶段子参数敏感性、真实精度效应(fp16 vs INT8)、以及跨实验偏倚——后者对不同模型产生相反扭曲,故两篇前期情绪向量研究间的单一rho值不可直接解释。
原文摘要 · Abstract (English)
We extract 21-emotion vector sets from twelve small language models (six architectures x base/instruct, 1B-8B parameters) under a unified comprehension-mode pipeline at fp16 precision, and compare the resulting geometries via representational similarity analysis on raw cosine RDMs. The five mature architectures (Qwen 2.5 1.5B, SmolLM2 1.7B, Llama 3.2 3B, Mistral 7B v0.3, Llama 3.1 8B) share nearly identical 21-emotion geometry, with pairwise RDM Spearman correlations of 0.74-0.92. This universality persists across diametrically opposed behavioral profiles: Qwen 2.5 and Llama 3.2 occupy opposite poles of MTI Compliance facets yet produce nearly identical emotion RDMs (rho = 0.81), so behavioral facet differences arise above the shared emotion representation. Gemma-3 1B base, the one immature case in our dataset, exhibits extreme residual-stream anisotropy (0.997) and is restructured by RLHF across all geometric descriptors, whereas the five already-mature families show within-family base x instruct RDM correlations of rho >= 0.92 (Mistral 7B v0.3 at rho = 0.985), suggesting RLHF restructures only representations that are not yet organized. Methodologically, we show that what prior work has read as a single comprehension-vs-generation method effect in fact decomposes into four distinct layers -- a coarse method-dependent dissociation, robust sub-parameter sensitivity within generation, a true precision (fp16 vs INT8) effect, and a conflated cross-experiment bias that distorts in opposite directions for different models -- so that a single rho between two prior emotion-vector studies is not a safe basis for interpretation without the layered decomposition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。