arXiv:2510.22042cs.CLcs.AI2025-10被引 13

发现大模型内部存在可操控的情感空间,跨语言通用且稳定。

Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models

  • 通过分析隐藏状态几何结构,发现情感在低维流形上定向分布。
  • 跨域对齐后线性探测误差低,八种语言数据集表现一致。
  • 可学习干预模块精准控制基本情绪,保持语义不变。

本研究探讨大语言模型(LLMs)如何在内部表征情感,通过分析其隐藏状态空间的几何结构,发现存在一个低维情感流形。情感表示以方向性方式编码,分布在各层中,并与可解释维度对齐。该结构在深度上保持稳定,且在涵盖五种语言的八个真实世界情感数据集上具有泛化能力。跨域对齐后表现出低误差和强线性探测性能,表明存在一个通用情感子空间。在该空间内,可通过学习的干预模块实现内部情感感知的调控,同时保持语义完整性,尤其在基本情绪上的控制效果显著。这些发现揭示了大模型中一致且可操控的情感几何结构,为理解其情感内化与处理机制提供了新视角。

原文摘要 · Abstract (English)

This work investigates how large language models (LLMs) internally represent emotion by analyzing the geometry of their hidden-state space. The paper identifies a low-dimensional emotional manifold and shows that emotional representations are directionally encoded, distributed across layers, and aligned with interpretable dimensions. These structures are stable across depth and generalize to eight real-world emotion datasets spanning five languages. Cross-domain alignment yields low error and strong linear probe performance, indicating a universal emotional subspace. Within this space, internal emotion perception can be steered while preserving semantics using a learned intervention module, with especially strong control for basic emotions across languages. These findings reveal a consistent and manipulable affective geometry in LLMs and offer insight into how they internalize and process emotion.

情感建模语言模型隐空间分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。