arXiv:2604.07382cs.LGcs.AI2026-04被引 4

发现大模型情感表征具有类心理学的几何结构。

Latent Structure of Affective Representations in Large Language Models

  • 用几何分析法研究大模型中的情感表征空间。
  • 表征与心理学的唤醒-效价模型高度吻合,且可线性近似。
  • 可据此量化情感任务中的模型不确定性,提升可解释性。

大型语言模型(LLMs)潜在表征的几何结构是当前研究热点,尤其关乎模型透明度与人工智能安全。现有研究多关注表征的一般几何与拓扑特性,但因缺乏真实潜空间几何作为基准,验证困难。情绪处理为探测表征几何提供了理想场景,因其兼具分类组织与连续效价维度,且在心理学中已有充分验证。本文利用几何数据分析工具,探究了LLMs中情感表征的潜在结构。主要发现包括:第一,LLMs学习到的潜空间情感表征与心理学广泛采用的唤醒-效价模型一致;第二,这些表征呈现非线性几何结构,但仍可被良好线性近似,支持了模型透明性方法中常见的线性假设;第三,所学潜空间可用于量化情感任务中的不确定性。结果表明,LLMs获得的情感表征具备与人类情绪模型平行的几何结构,对模型可解释性与安全性具有实际意义。

原文摘要 · Abstract (English)

The geometric structure of latent representations in large language models (LLMs) is an active area of research, driven in part by its implications for model transparency and AI safety. Existing literature has focused mainly on general geometric and topological properties of the learnt representations, but due to a lack of ground-truth latent geometry, validating the findings of such approaches is challenging. Emotion processing provides an intriguing testbed for probing representational geometry, as emotions exhibit both categorical organization and continuous affective dimensions, which are well-established in the psychology literature. Moreover, understanding such representations carries safety relevance. In this work, we investigate the latent structure of affective representations in LLMs using geometric data analysis tools. We present three main findings. First, we show that LLMs learn coherent latent representations of affective emotions that align with widely used valence--arousal models from psychology. Second, we find that these representations exhibit nonlinear geometric structure that can nonetheless be well-approximated linearly, providing empirical support for the linear representation hypothesis commonly assumed in model transparency methods. Third, we demonstrate that the learned latent representation space can be leveraged to quantify uncertainty in emotion processing tasks. Our findings suggest that LLMs acquire affective representations with geometric structure paralleling established models of human emotion, with practical implications for model interpretability and safety.

情感表征几何结构模型可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。