语音语言模型隐式完成语音转写,提升语音理解能力。
Interleaved Speech Language Models Latently Work In Text

- 通过日志透镜分析发现模型在中间层隐式转写语音为文本。
- 77%数据中,正确文本词成为最高候选,无需专门训练。
- 适合研究多模态模型机制与语音理解优化的研究者。
语音语言模型(SLMs)广泛研究,主流方法是将语音与文本数据交错训练,以增强纯语音能力。然而,两种模态在模型隐空间中的交互机制尚不明确。本文通过日志透镜分析不同规模和架构的交错语音-文本模型,揭示模型在推理过程中经历一个隐式转写阶段:尽管未针对语音识别训练,但语音对应的文本词在中间层已可解码。该转写结果在高达77%的数据中位列候选首位。随后模型在文本空间预测下一个词,再转换回语音域。我们进一步分析了交错数据和从文本模型初始化对这一行为的影响,并考察其与语音知识能力的相关性。研究揭示了语音与文本模态间的内在机制,为优化语音语言模型提供新思路。
原文摘要 · Abstract (English)
Speech language models (SLMs) have been extensively studied, with the common paradigm incorporating text data and pre-trained text LMs. A leading approach is speech-text interleaving in which models are trained over sequences containing both speech and text tokens, aiming to boost even speech-only capabilities. Yet the way these two modalities interact in the model latent space remains unclear. In this work, we analyze interleaved speech-text LMs from different model families and sizes through the scope of the logit lens to provide such insight. We reveal that these models go through an implicit transcription phase in which the text token of the spoken word becomes decodable in intermediate layers, despite not being trained for speech recognition. The transcription of the word appears as one of the top candidate words for as much as 77\% of the data. Following this stage, the models proceed to predict the next word in the text space before transforming back to the speech domain. We finally analyze the role of interleaving data, and initializing from text LMs in eliciting this behavior, as well as seeing how this correlates with spoken knowledge abilities. Our analysis sheds light on the internal mechanisms underlying the relationship between speech and text modalities and could shape SLM optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。