为大模型情感智能评估与提升提供心理学框架和实证工具
EiCAP: Beyond Fluency, Probing and Improving Emotional Intelligence in LLMs via Psychologically Grounded Multi-Turn Dialogue
- 构建六层心理基础的情感智能分类体系,配套测评与训练数据集
- 基于该体系,微调后模型在24个子任务上平均得分达75.33%
- 直接情感智能训练优于通用对话微调,且仅需0.8%参数量
大型语言模型越来越多地应用于心理健康支持、教育和危机应对等情感敏感场景,但缺乏系统评估和提升情感智能(EI)的框架。本文提出EiCAP,一个基于心理学的六层情感智能分类体系,包含两个互补资源:EiCAP-Bench是一个包含3,174个测试样本的多轮强制选择测评集,涵盖24个子类别及跨轮次依赖关系;EiCAP-SFT是一个152,820条对话的监督语料库,与同一分类体系对齐。关键发现:通用对话微调(如UltraChat)无法提升情感智能,在24个子类别中平均得分为24.6%,接近随机水平(25%);而采用仅占0.8%参数量的基于情感智能的LoRA微调,使Qwen-2.5-7B-Base模型在所有24个子类别上平均得分达75.33%,较基础模型提升51.7个百分点,较指令模型提升37.1个百分点。消融实验表明,前置的UltraChat预训练反而降低性能21.4个百分点,证明直接情感智能训练既必要又充分。
原文摘要 · Abstract (English)
Large Language Models increasingly serve in emotionally sensitive roles, including mental health support, education, and crisis response, yet they lack a principled framework for assessing or improving Emotional Intelligence (EI). We introduce EiCAP, a unified, psychologically grounded six-layer EI taxonomy operationalized into two complementary resources. EiCAP-Bench is a multi-turn, one-vs-three forced-choice evaluation suite with 3,174 probes across 24 subcategories and cross-turn dependencies that reflect real conversational EI demands. EiCAP-SFT is a 152,820-dialogue supervision corpus aligned to the same taxonomy, enabling controlled, interpretable fine-tuning. Two key findings emerge. First, generic conversational supervised fine-tuning does not confer EI: fine-tuning on UltraChat yields no significant gain in any of the 24 subcategories, with a macro score of 24.6%, near the chance level of 25%. Second, applying EI-grounded LoRA, using approximately 0.8% of parameters, directly to Qwen-2.5-7B-Base achieves significant gains in all 24 subcategories, reaching a macro score of 75.33%, a gain of 51.7 percentage points over Base and 37.1 percentage points over Instruct. Crucially, an ablation shows that the UltraChat pre-stage is counterproductive, reducing performance by 21.4 percentage points: direct EI-grounded training is both necessary and sufficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。