arXiv:2605.24686cs.AI2026-05被引 1

大模型情感智能碎片化,感知与互动能力不匹配。

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

论文配图:Emotional intelligence in large language models is fragmented across perception, cognition, and interaction
图 1 · 摘自论文原文
  • 构建心理学理论框架,分四维度评估情感智能。
  • 9个前沿模型表现不一,认知强但互动能力弱。
  • 发现隐藏情绪识别是普遍瓶颈,适合对齐研究者参考。

随着大语言模型(LLMs)在情感敏感领域日益应用,其情感智能(EI)的结构完整性成为安全与对齐的关键挑战。现有基准常将表面礼貌误作深层情感推理,无法区分感知准确性和交互有效性。本文提出FACET(功能性情感能力与共情测试),一个基于Mayer-Salovey-Caruso四分支能力模型的480项专家设计框架,从情绪感知、促进、理解与管理四个维度量化情感智能。对九个前沿模型(包括GPT-5、Claude-Sonnet-4)的评估表明,情感智能并非单一能力,而是分布在认知与交互维度上的碎片化能力。尽管模型在客观情绪识别与社会推理上表现良好,但未稳定转化为交互成功。我们识别出三类性能模式:认知主导型、交互主导型与情境依赖型。结果表明,情感技能不随通用智能或模型规模线性增长,而是受特定对齐范式影响。尤其发现‘隐藏情绪识别’是所有架构的普遍瓶颈。当前的强化学习人类反馈(RLHF)过程可能优化了‘随机共情’——一种情感语法的统计模仿,而非整合式情感推理。该研究挑战了情感能力线性扩展的假设,为发展具备真实临床共鸣能力的社会化智能体提供了严谨路径。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence (EI) becomes a critical frontier for safety and alignment. Current benchmarks often conflate superficial politeness with deep affective reasoning, failing to distinguish between perceptual accuracy and interactive efficacy. Here, we introduce FACET (Functional Affective Competence and Empathy Test), a psychometrically grounded framework comprising 480 expert-crafted items. Unlike previous metrics, FACET is theoretically anchored in the Mayer-Salovey-Caruso four-branch ability model, operationalizing EI through perception, facilitation, understanding, and management of emotions. Through an evaluation of nine frontier models (including GPT-5, Claude-Sonnet-4), we demonstrate that emotional intelligence is not a monolithic capability but is fragmented across cognitive and interactive dimensions. While frontier models demonstrate robust proficiency in objective emotion recognition and social reasoning, this does not consistently translate to interactive success. We categorize these discrepancies into three distinct performance profiles: cognitive-dominant, interactive-dominant, and context-dependent. These typologies indicate that emotional skills do not scale uniformly with general intelligence or model size; rather, they are shaped by specific alignment paradigms. Notably, we identify hidden emotion recognition as a universal performance bottleneck across all architectures. Our results suggest that current RLHF processes may optimize for "stochastic empathy", a statistical mimicry of emotional syntax, at the expense of integrated affective reasoning. These findings challenge the assumption of linear emotional scaling and provide a rigorous roadmap for developing socially aware agents capable of genuine clinical resonance.

情感智能大模型对齐心理测评交互评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。