arXiv:2508.08095cs.CLcs.AI2025-08

让语音语言模型更好理解情绪,通过解耦语音与语言信息提升对话表现。

Dual Information Speech Language Models for Emotional Conversations

  • 设计异构适配器解耦语音和语言信息,避免相互干扰。
  • 弱监督训练下在情感对话任务中达到可比性能,仅训练适配器。
  • 适合需要融合语音情绪与语义的对话系统开发者使用。

依赖文本大模型的对话系统常忽略语气等副语言线索,而语音-语言模型(SLMs)以语音为输入,正成为解决此问题的可行方案。然而,基于冻结大模型扩展的SLMs难以捕捉副语言信息且上下文理解能力下降。本文指出信息纠缠和不当训练策略是核心问题。为此,提出两种异构适配器并设计弱监督训练策略,实现副语言与语言信息的解耦,使模型通过结构化表示理解语音,同时通过受控随机性避免生成特定任务向量,保持上下文理解。该方法仅在通用数据集上训练适配器,兼顾参数与数据效率。实验表明,模型在情感对话任务中表现优异,有效融合了副语言与语言信息于上下文之中。

原文摘要 · Abstract (English)

Conversational systems relying on text-based large language models (LLMs) often overlook paralinguistic cues, essential for understanding emotions and intentions. Speech-language models (SLMs), which use speech as input, are emerging as a promising solution. However, SLMs built by extending frozen LLMs struggle to capture paralinguistic information and exhibit reduced context understanding. We identify entangled information and improper training strategies as key issues. To address these issues, we propose two heterogeneous adapters and suggest a weakly supervised training strategy. Our approach disentangles paralinguistic and linguistic information, enabling SLMs to interpret speech through structured representations. It also preserves contextual understanding by avoiding the generation of task-specific vectors through controlled randomness. This approach trains only the adapters on common datasets, ensuring parameter and data efficiency. Experiments demonstrate competitive performance in emotional conversation tasks, showcasing the model's ability to effectively integrate both paralinguistic and linguistic information within contextual settings.

语音理解情感对话适配器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。