探索大模型是否有意识,提出三种可能的身份认定方式。
Where is the Mind? Persona Vectors and LLM Individuation

- 从注意力流分析出发,认为模型存在虚拟实例身份。
- 发现人格向量在模型中形成独立空间,支持人格化身份。
- 适合关注模型本质与意识的哲学与AI研究者阅读。
大语言模型的个体化问题在于:哪些相关实体可被视作具有心智。本文通过机制可解释性方法,结合近期关于人格向量、人格空间及涌现错位的实证研究,提出三种最强候选观点:虚拟实例观,以及本文提出的两种新观点——(虚拟)实例-人格观和模型-人格观。首先,基于注意力流在词元时间上维持类心理连接,支持虚拟实例观;其次,梳理人格研究,围绕三个关于模型内部人格结构的假设,证明基于人格的两种观点具有前景。
原文摘要 · Abstract (English)
The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds. We approach this problem through mechanistic interpretability, engaging in particular with recent empirical work on persona vectors, persona space, and emergent misalignment. We argue that three views are the strongest candidates: the virtual instance view and two new views we introduce, the (virtual) instance-persona view and the model-persona view. First, we argue for the virtual instance view on the grounds that attention streams sustain quasi-psychological connections across token-time. Then we present the persona literature, organised around three hypotheses about the internal structure underlying personas in LLMs, and show that the two persona-based views are promising alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。