发现大模型中可组合的文学特征,能精准控制情感与风格。
Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect
- 用稀疏自编码器分析模型中间层,识别出四类文学特征。
- Llama 在 27 类情绪中全覆盖,Gemma 缺少 4 类,主要在'钦佩'上失败。
- 适合研究模型可解释性、情感生成或提示工程的读者。
本文通过在两个指令微调的大语言模型(Llama 3.1 8B-Instruct 和 Gemma 2 9B-IT)的中深层残差流上使用稀疏自编码器,刻画了文学基本单元的组合架构。四类特征浮现:促进目标情感词汇的命名门控、第一人称语体的十一自体簇、风格调节器(表现而非讲述、陌生化),以及仅由多特征协同产生的复合情绪。在 27 类情绪分类(Cowen-Keltner)的五模型裁判评估中,Llama 达到 27/27 全覆盖,依赖命名门控、多特征配方和单自体特征操控;Gemma 达到 23/27,唯一未达标的是‘钦佩’。随机判断下,单细胞通过概率约为 $10^{-3}$,两种子出现的假阳性总数可忽略,表明观察到的覆盖度非偶然。跨架构差异在于严格与宽松裁判的一致性:相同生成结果中,裁判对 Llama 输出的共识高于 Gemma,因 Llama 更直接命名情感,而 Gemma 通过场景与意象唤起情感。两类模型均存在兼具语体标记与情感发射功能的自体特征,每架构有一个最强化的自体特征,在特定配置下增强机构型助手人格,并在同一校准系数下生成可归类的情感输出。方法上提出三阶段验证流程(逻辑透镜、LLM评分、五模型裁判),包含已知反模式,单次情绪特征发现周期仅需单卡约 15 分钟。
原文摘要 · Abstract (English)
We characterize a compositional architecture of literary primitives in two instruction-tuned large language models (Llama 3.1 8B-Instruct and Gemma 2 9B-IT) via sparse autoencoders on mid-depth residual streams. Four feature classes emerge: naming-gates that promote lexical tokens of a target affect, an eleven-self cluster of first-person register features, stylistic register modulators (show-don't-tell and defamiliarization), and compositional emotions that arise only from multi-feature steering. Under a forced-choice 5-LLM judge panel applied to a 27-category emotion taxonomy (Cowen-Keltner), Llama reaches full 27/27 coverage by combining naming-gates, multi-feature recipes, and single self-feature steering; Gemma reaches 23/27 with adoration as the single residual strict-fail. Under random judging, the per-cell pass probability is on the order of $10^{-3}$ and the expected number of two-seed false-positive cells across the catalog is negligible, so the observed coverage is not consistent with chance. A cross-architectural asymmetry sits in the strict-versus-soft judge contrast: on the same generations, judges agree more often on Llama outputs than on Gemma outputs because Llama outputs name the target affect more directly while Gemma outputs evoke it through scene and imagery. Both architectures contain self-features that serve simultaneously as register markers and as emotion emitters, including a single most-RLHF-loaded self-feature per architecture that intensifies the institutional Helper-AI persona at one operating regime and produces affect-categorizable output at the same calibrated coefficient. Methodologically, the paper presents a three-stage validation pipeline (logit-lens, LLM-rate, 5-LLM judge) with documented anti-patterns; the total compute is single-GPU and about 15 minutes per emotion-feature discovery cycle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。