arXiv:2605.29685cs.AI2026-05

首个基于社会理论的LLM社交智能诊断基准,可精准定位模型短板。

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

论文配图:NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs
图 1 · 摘自论文原文
  • 构建4大类11维度的社会智能框架,细化为可测能力点。
  • 在137个中文情境测试中,模型通信能力普遍薄弱,尤其三方面表现差。
  • 适合评估客服、陪伴类AI的社交缺陷,推动安全应用落地。

随着大语言模型在情感陪伴、客户服务等社会场景中的广泛应用,其社交智能的评测对人机交互质量与安全至关重要。现有基准缺乏统一框架,难以实现细粒度诊断。为此,我们通过文献综述与多阶段专家验证,构建首个基于心理测量学原则的社会智能框架,包含4个类别、11个维度及细粒度能力分支。基于此框架,提出NICE(Norm, Interaction, Cognition, Experience)诊断基准,涵盖137个代表性中文情境。在5个前沿大模型与人类参照组的测试中,模型整体准确率较高,但通信能力存在系统性弱点,框架将其定位至三个具体能力点:多轮对话、非语言沟通和同步性。NICE将社交智能评估转向理论驱动的精细化诊断,揭示模型在社会关键环节的薄弱之处。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social intelligence has become critical to the quality and safety of human-AI interaction. However, existing social intelligence benchmarks lack a unified framework that organizes social abilities into a unified structure, and therefore cannot enable fine-grained diagnosis. To build the first holistic diagnostic evaluation grounded in social theory, we first construct a social intelligence framework through a literature review and multi-stage expert validation guided by psychometric principles. The resulting framework includes 4 categories and 11 dimensions, each further specified by fine-grained capability facets. Building on this framework, we introduce NICE (Norm, Interaction, Cognition, Experience), a diagnostic benchmark of 137 items operationalized through representative Chinese contexts. Across 5 frontier LLMs and a human reference group, models score higher in aggregate accuracy yet show a consistent weakness in Communication, which the framework localizes to 3 specific capability facets: multi-turn communication, nonverbal communication, and synchrony. NICE thus reframes social intelligence evaluation toward theory-grounded diagnosis of socially consequential weaknesses in LLMs.

社交智能评测基准大模型评估理论驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。