多轮对话中语音模型会遗忘指定说话风格,导致表达失真。
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
- 通过显式提醒可缓解语音风格遗忘问题
- 系统消息中的风格指令效果差于用户消息
- 情绪、语速等副语言风格难以持续保持
本文揭示了在多轮对话中,当语音语言模型(SLMs)被要求以特定说话风格开场时,无法在后续回合中持续保持该风格,我们称之为语音模型的风格遗忘现象。研究聚焦于情感、口音、音量和语速等副语言风格。评估了三种专有和两种开源的SLMs,结果表明所有模型均无法维持指定风格。尽管模型能回忆起风格指令,但实际表达仍失败;通过显式提示可缓解此问题。此外,风格指令置于系统消息中时表现更差,即便系统消息专为持久性指令设计。研究揭示了当前SLMs在风格保持上的系统性缺陷,强调未来模型需提升风格一致性。代码与数据已公开于https://github.com/YuXiangLin1234/SLM-Style-Amnesia。
原文摘要 · Abstract (English)
In this paper, we show that when spoken language models (SLMs) are instructed to speak in a specific speaking style at the beginning of a multi-turn conversation, they cannot maintain the required speaking styles after several turns of interaction; we refer to this as the style amnesia of SLMs. We focus on paralinguistic speaking styles, including emotion, accent, volume, and speaking speed. We evaluate three proprietary and two open-source SLMs, demonstrating that none of these models can maintain a consistent speaking style when instructed to do so. We further show that while SLMs can recall the style instruction when prompted in later turns, they still fail to express it, but through explicit recall can mitigate style amnesia. In addition, SLMs struggle more when the style instruction is placed in system messages rather than user messages, even though system messages are specifically designed to provide persistent, conversation-level instructions. Our findings highlight a systematic gap in current SLMs' ability to maintain speaking styles, highlighting the need for improved style adherence in future models. Our code and evaluation data are publicly available at https://github.com/YuXiangLin1234/SLM-Style-Amnesia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。