让语音模型学会理解情绪来源,生成更像人的回应。
A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction
- 将用户情绪及成因注入模型内部推理,实现情感内化
- 在HumDial基准上三项指标均排名第一,优于大模型与人工评估
- 适合需要高情商交互的对话系统开发者
本文提出一种统一的语音语言模型,通过新型数据构建策略——注入情绪归因思维(IEAT),将用户情绪状态及其潜在原因融入模型内部推理过程,使情感感知成为内在推理而非显式监督。模型采用两阶段渐进式训练:第一阶段通过自蒸馏完成语音-文本对齐与情绪属性建模;第二阶段进行端到端跨模态联合优化,确保文本与语音情感表达一致。在人类类语音对话系统挑战赛(HumDial)情感智能基准上的实验表明,该方法在情绪轨迹建模、情感推理与共情响应生成三项任务中,均在基于大模型与人工评估下取得领先性能。
原文摘要 · Abstract (English)
This paper presents a unified spoken language model for emotional intelligence, enhanced by a novel data construction strategy termed Injected Emotional-Attribution Thinking (IEAT). IEAT incorporates user emotional states and their underlying causes into the model's internal reasoning process, enabling emotion-aware reasoning to be internalized rather than treated as explicit supervision. The model is trained with a two-stage progressive strategy. The first stage performs speech-text alignment and emotional attribute modeling via self-distillation, while the second stage conducts end-to-end cross-modal joint optimization to ensure consistency between textual and spoken emotional expressions. Experiments on the Human-like Spoken Dialogue Systems Challenge (HumDial) Emotional Intelligence benchmark demonstrate that the proposed approach achieves top-ranked performance across emotional trajectory modeling, emotional reasoning, and empathetic response generation under both LLM-based and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。