arXiv:2606.16843cs.CL2026-06

用Transformer模型验证情绪空间是否符合心理学理论。

Data-Driven Decoding of Russell's Circumplex Model of Affect

论文配图:Data-Driven Decoding of Russell's Circumplex Model of Affect
图 1 · 摘自论文原文
  • 融合文本与语音的Transformer模型可还原情绪拓扑结构。
  • 多模态融合使情绪坐标与人类标注完全对齐。
  • 无需微调即可在通用文本嵌入中定位精细情绪词。

情感计算日益依赖深度学习表示情绪,但潜在空间常被视为高维黑箱。本文检验Transformer嵌入能否恢复Russell情绪环状模型的几何规律。通过在自然语境数据集MSP-Podcast和受控的LLM生成刺激上训练文本(RoBERTa)与语音(wav2vec 2.0)编码器,并构建多模态Transformer融合架构,分析其潜在空间是否符合效价-唤醒度维度及人类感知邻近关系。结果表明,文本与音频的多模态融合实现了与Russell模型完全一致的拓扑对齐;在零样本设置下,通用文本嵌入投影出的细粒度情绪词接近其已知的人类映射坐标。本研究提出一种数据驱动的情绪模型验证框架,证明情绪环状结构并非仅源于人为标注,而是内在编码于多种模态的表示中,弥合了心理理论与表征学习之间的鸿沟。

原文摘要 · Abstract (English)

Affective computing increasingly relies on deep learning to represent emotions, yet latent spaces often remain opaque, high-dimensional black boxes. This paper investigates whether Transformers' embeddings recover the geometric regularities of Russell's circumplex model. We unify two complementary experiments testing the hypothesis that, after training models on text and speech, their resulting latent spaces encode a topology consistent with valence-arousal and reproduce human-like neighborhood relations. Specifically, we evaluate deep representations extracted from Transformer-based text (RoBERTa) and speech (wav2vec 2.0) encoders, along with a multimodal Transformer fusion architecture, across naturalistic datasets like MSP-Podcast and controlled LLM-generated stimuli. Our analysis reveals that multimodal fusion of text and audio yields perfect topological alignment with Russell's primary emotion ordering. Furthermore, in a zero-shot setting using generic text embeddings, projected fine-grained emotion terms fall close to their established human-mapped coordinates. Our contribution is a novel, data-driven framework for validating emotion models, demonstrating that Russell's circumplex structure is intrinsically encoded in the embeddings of these modalities rather than being solely an artifact of human labeling, thereby bridging the gap between psychological theory and representation learning.

情绪识别多模态情感建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。