研究语音模型如何通过上下文学习,发现语速关键而音高影响小。
In-Context Learning in Speech Language Models: Analyzing the Role of Acoustic Features, Linguistic Structure, and Induction Heads
- 分析语音模型在上下文学习中对声学与语言特征的依赖
- 语速显著影响学习效果且被输出模仿,音高则无明显作用
- 发现归纳头是语音上下文学习的关键机制,可完全消除学习能力
在仅文本语言模型中,上下文学习(ICL)已得到广泛研究,但在语音领域仍基本未被探索。本文研究语言和声学特征如何影响语音语言模型中的上下文学习。聚焦于文本转语音(TTS)任务,从两个角度分析:(1) 模型能否准确从示例中推断出任务(即生成正确的语音内容),以及 (2) 模型输出是否模仿示例语音的声学特征。结果表明,语速对ICL性能有显著影响,并在输出中被一致模仿;而音高范围和音量对性能影响较小,且不一致地再现。最后,我们研究了归纳头在语音上下文学习中的作用,发现这些头部具有因果作用:移除前k个归纳头会完全消除模型的ICL能力,与纯文本模型中的发现一致。
原文摘要 · Abstract (English)
In-Context Learning (ICL) has been extensively studied in text-only Language Models, but remains largely unexplored in the speech domain. Here, we investigate how linguistic and acoustic features affect ICL in Speech Language Models. We focus on the Text-to-Speech (TTS) task, which allows us to analyze ICL from two angles: (1) how accurately the model infers the task from the demonstrations (i.e., generating the correct spoken content), and (2) to what extent the model mimics the acoustic characteristics of the demonstration speech in its output. We find that speaking rate strongly affects ICL performance and is also mimicked in the output, whereas pitch range and intensity have little impact on performance and are not consistently reproduced. Finally, we investigate the role of induction heads in speech-based ICL and show that these heads play a causal role: ablating the top-k induction heads completely removes the model's ICL ability, mirroring findings from text-based ICL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。