让语音大模型学会共情,精准回应情绪。
Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models

- 构建去语义主导的多情绪数据集,分离声学与语义特征。
- 引入心理推理链,生成有逻辑且富有同理心的回复。
- 动态平衡推理严谨性与情感响应质量,适合对话系统优化。
尽管语音语言模型(ALMs)具备强大的语义理解能力,但在复杂情感交互中仍表现不佳。文本语义主导常掩盖声学细节,认知深度不足导致回复泛化、缺乏情绪感知。我们提出CogAudio-LLM框架,通过构建LIME-440K数据集(语义相同但情绪多样),实现声学与语义解耦。引入四步心理推理链(EIPS),并通过分阶段训练将该推理过程显式微调后隐式蒸馏至生成流程。最后设计双路径软自适应策略优化(DR-SAPO),动态调节推理逻辑与直接回应的情感质量,显著提升模型在情感交互中的表现。
原文摘要 · Abstract (English)
While Audio Language Models (ALMs) demonstrate strong semantic understanding, they struggle with complex affective interactions. Specifically, textual semantic dominance often overshadows acoustic nuances, and a lack of cognitive depth leads to generic, emotion-agnostic responses. We propose CogAudio-LLM\footnote{ \urlstyle{same} https://github.com/zxzhao0/CogAudio-LLM, a novel cognitive affective reasoning framework. To mitigate semantic dominance, we build LIME-440K, a ``lexically-identical, multi-emotion'' dataset designed to facilitate acoustic-semantic decoupling. We introduce EIPS, a 4-step Chain-of-Thought (CoT) mechanism incorporating psychological reasoning. For inference efficiency, multi-stage training explicitly establishes EIPS via supervised fine-tuning, then distills this logic into an implicit generation process. Finally, we design DR-SAPO (Dual-Route Soft Adaptive Policy Optimization) to dynamically balance the logical rigor of the CoT with the empathetic quality of the direct response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。