剖析语音大模型生成不连贯的三大原因,发现语调影响最严重。
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
- 逐步从文本转为语音,分离验证三类因素的影响
- 语调复杂性对词汇建模影响最大,序列长度次之
- 为端到端语音模型训练提供关键优化方向
尽管基于文本的大语言模型已具备人类水平的写作能力与显著智能,语音语言模型(SLMs)仍难以生成语义连贯的输出。可能的原因包括:(A) 语音标记主要传递音素信息而非语义信息,(B) 语音序列长度远超文本序列,(C) 非语言信息(如语调)引入额外复杂性和变异性。本文通过逐步从文本向语音演进的模态转换方式,分别探究这三个关键因素的影响。研究发现,因素A影响相对较小,因素B对句法与语义建模影响更明显,而因素C影响最为显著,尤其在基础词汇建模层面。基于此,论文揭示了训练SLMs的独特挑战,并指明提升端到端语音模型有效性的路径。
原文摘要 · Abstract (English)
Although text-based large language models exhibit human-level writing ability and remarkable intelligence, speech language models (SLMs) still struggle to generate semantically coherent outputs. There are several potential reasons for this performance degradation: (A) speech tokens mainly provide phonetic information rather than semantic information, (B) the length of speech sequences is much longer than that of text sequences, and (C) paralinguistic information, such as prosody, introduces additional complexity and variability. In this paper, we explore the influence of three key factors separately by transiting the modality from text to speech in an evolving manner. Our findings reveal that the impact of the three factors varies. Factor A has a relatively minor impact, factor B influences syntactical and semantic modeling more obviously, and factor C exerts the most significant impact, particularly in the basic lexical modeling. Based on these findings, we provide insights into the unique challenges of training SLMs and highlight pathways to develop more effective end-to-end SLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。