arXiv:2510.02569cs.CL2025-10被引 3

探究语音模型中间表示,发现其用英语语义或发音表达跨语言信息。

Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models

  • 通过寻找最接近的解码器词,分析语音模型中间表示机制。
  • 使用Whisper编码器的模型采用英语语义中间表征,可处理未见语言。
  • 非Whisper模型则用英语词汇表达语音音素,依赖训练目标差异。

集成语音与大语言模型的语音语言模型(SLMs)依赖模态适配器(MAs)将语音编码器输出映射为解码器语言模型可理解的表示。然而我们对这些关键适配器如何转换表示知之甚少。本文研究了三种SLMs(SALMONN、Qwen2-Audio和Phi-4-Multimodal-Instruct)中MA的输出表示。通过寻找与MA表示最近的解码器词,揭示出两种策略:使用Whisper编码器的模型倾向于以基于英语的中间语义(interlingua)表征输入含义,从而支持指令微调中未出现的语言;而未使用Whisper的模型(如Phi-4-Multimodal-Instruct)则以英语词汇表达输入的语音特征。我们推测这取决于语音编码器是否仅训练用于语音识别,还是也包含翻译任务。

原文摘要 · Abstract (English)

Spoken language models (SLMs) that integrate speech with large language models (LMs) rely on modality adapters (MAs) to map the output of speech encoders to a representation that is understandable to the decoder LM. Yet we know very little about how these crucial MAs transform representations. Here we examine the MA output representation in three SLMs (SALMONN, Qwen2-Audio and Phi-4-Multimodal-Instruct). By finding the nearest decoder LM token to an MA representation, we uncover two strategies for MA representations. For models using a Whisper encoder, MAs appear to represent the meaning of the input using an English-based interlingua, allowing them to handle languages unseen in instruction tuning. For models that don't, like Phi-4-Multimodal-Instruct, MAs instead represent the phonetics of the input, but expressed with English words. We hypothesise that which arises depends on whether the speech encoder is trained only for speech recognition or also for translation.

语音模型中间表征跨语言Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。