用冻结大模型+语音数据训练情感对话机器人,高效又不失表达力。
FreezeEmpath: Efficient Training for Empathetic Spoken Chatbots with Frozen LLMs

- 冻结大模型参数,仅用语音指令与情绪识别数据训练
- 在情感对话、语音情绪识别等任务上超越现有模型
- 适合需要快速部署情感语音交互系统的团队
情感是实现自然语音对话的关键,使机器能识别人类言语中的情绪并作出共情回应。近年来,基于大语言模型(LLMs)的共情语音聊天机器人取得显著进展。然而,训练仍面临挑战:依赖昂贵的情感语音指令数据,且生成语音缺乏情绪表现力;使用跨模态情感指令微调可能导致灾难性遗忘,削弱模型通用能力。为此,我们提出 FreezeEmpath,一种端到端的共情语音聊天机器人,训练方式简单高效。整个过程仅使用现有语音指令数据和语音情绪识别(SER)数据,保持 LLM 参数冻结。实验表明,FreezeEmpath 能生成富有情感表达的语音,在共情对话、语音情绪识别(SER)及语音问答(SpokenQA)任务中均优于其他模型,验证了该训练策略的有效性。
原文摘要 · Abstract (English)
Empathy is essential for fostering natural interactions in spoken dialogue systems, as it enables machines to recognize the emotional tone of human speech and deliver empathetic responses. Recent research has made significant progress in developing empathetic spoken chatbots based on large language models (LLMs). However, several challenges still exist when training such models, including reliance on costly empathetic speech instruction data and a lack of emotional expressiveness in the generated speech. Finetuning LLM with cross-modal empathetic instruction data may also lead to catastrophic forgetting and a degradation of its general capability. To address these challenges, we propose FreezeEmpath, an end-to-end empathetic spoken chatbot trained in a simple and efficient manner. The entire training process relies solely on existing speech instruction data and speech emotion recognition (SER) data, while keeping the LLM's parameters frozen. Experiments demonstrate that FreezeEmpath is able to generate emotionally expressive speech and outperforms other empathetic models in empathetic dialogue, SER, and SpokenQA tasks, demonstrating the effectiveness of our training strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。