arXiv:2608.07852cs.CL2026-08

揭示大模型中助手与角色扮演人格的内在关联与演变机制

"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

论文配图:"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders
图 1 · 摘自论文原文
  • 用稀疏自编码器分析模型在不同对话模式下的内部表示
  • 角色扮演人格保留助手核心特征,逐层演化出行为与风格差异
  • 故事角色无助手核心特征,但默认模式下助手可能缓慢进入角色

语言模型如何内化说话者身份——如助手、角色扮演人格或故事叙述角色——仍不明确。本文基于用户表达的情绪文本及对应模型回复,研究三种生成场景(助手、角色扮演、故事)中的说话者表征。通过在回合边界和代词标记位置提取稀疏自编码器特征,并经筛选管道选择不同深度的特征,分析其控制效应与激活分布。主要发现:助手与角色扮演人格并非独立替代关系;角色人格保留助手相关的核心特征,并从操作机制层逐步向行为与风格层分化。而生成的故事角色缺乏该核心特征。尽管故事与角色扮演可通过沉浸式模拟模式区分,但在默认设置下,助手仍可能进入或缓慢漂移至角色状态。

原文摘要 · Abstract (English)

How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker representations using a dataset of user-expressed emotional text and corresponding model responses. We decompose three generation settings (Assistant, Roleplay, and Story) into sparse autoencoder features extracted at turn-boundary and pronoun-token positions and selected through a filtering pipeline for different depths. We characterize each surviving feature through its steering effects and activation distribution. Our main finding is that the Assistant and roleplay personas are not independent alternatives: personas retain the Assistant-associated feature core while progressively differentiating from it across layers, starting from operational machinery towards behavioral and stylistic features. Meanwhile, generated story characters lack the Assistant-associated core. Both Story and Roleplay can be distinguished from the Assistant with Immersive Simulation Mode. However, the Assistant can sometimes enter or slowly drift into it even in the default setting.

大模型推理角色表征自编码器语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。