arXiv:2505.19606cs.CL2025-05被引 4

语音编码器在跨语言对齐中既懂发音也懂语义,关键在语义而非发音。

Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically

  • 通过控制发音相似性,验证跨语言对齐源于语义而非发音。
  • 无发音线索时,翻译检索仍显著高于随机水平,尤其在额外训练翻译任务的模型中。
  • 早期退出编码器可提升低资源语言语音识别性能,适合多语言语音研究者。

预训练语言模型中的跨语言对齐可实现知识迁移。类似现象在Whisper风格的语音编码器中也被报道,基于语音翻译检索的表征相似性。然而,以往工作未控制等价语句间的发音重叠,可能人为支持检索结果。我们设计了发音控制实验,检验跨语言对齐是否源于语义而非发音相似性。结果显示,在以语音翻译为目标训练的编码器最终层中,即使去除发音线索,语音翻译检索性能仍远超随机水平,且在额外训练翻译任务的模型中表现尤为明显。此外,我们测试了提前退出编码器,以诱导更少受特定语言语义束缚的表示。实验表明,这能有效提升未参与训练的低资源语言上的自动语音识别性能。

原文摘要 · Abstract (English)

Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational similarity. However, prior work does not control for phonetic overlap between equivalent utterances, which may artificially support retrieval. We conduct pronunciation-controlled experiments to test whether cross-lingual alignment arises from semantic rather than phonetic similarity. Results show that spoken translation retrieval remains strongly above chance without phonetic cues in the final layers of encoders trained with a speech translation objective, most clearly for models additionally trained on translation. We further test early-exiting the encoder to induce representations we hypothesize to be less tied to language-specific semantics. These experiments indeed reveal performance gains in automatic speech recognition on low-resource languages unseen during training.

语音编码跨语言对齐低资源语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。