arXiv:2505.16182cs.SDeess.AS2025-05中稿 · Interspeech2025被引 2

用母语数据训练的离散语音单元,能提升非母语口音识别能力。

Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data

  • 基于自监督模型提取的离散语音单元模拟人类感知,实现无非母语数据的鲁棒识别。
  • 实验发现:共享母语的非母语听者比母语者更易理解目标口音语音,验证了跨语言可懂度优势。
  • 仅需母语数据即可提升对多种口音的适应性,适合资源稀缺语言的语音识别场景。

本研究揭示了一种仅使用母语语音数据实现抗口音语音识别的新思路。在人类对非母语语音的理解中,存在‘跨语言语音可懂度优势’(ISIB)现象:与说话人同源语言的非母语听者,理解效果甚至优于母语听者。基于自监督学习(SSL)模型提取的离散语音单元可反映人类对语音的感知,我们通过分析不同训练语言下离散语音单元在非母语语音上的表现,验证了该机制的技术实现。结果表明,离散语音单元确实表现出ISIB效应。由于方法仅依赖母语数据模拟人类感知行为,未来可广泛应用于语音数据稀缺的多种口音场景。

原文摘要 · Abstract (English)

In this study, we gained insight that contributes to achieving accent-robust ASR using only native speech data. In human perception of non-native speech, the phenomenon known as "interlanguage speech intelligibility benefit" (ISIB) is observed, where non-native listeners who share the native language with the speaker understand the speech better compared even to native listeners. Based on the idea that discrete tokens extracted from self-supervised learning (SSL) models represent the human perception of speech, we conducted an analytical study on the robustness of discrete token-based ASR to non-native speech, varying the language used for training the tokenization, which is viewed as a technical implementation of ISIB. The results showed that ISIB actually occurred in the discrete token-based ASR. Since our approach relies only on native speech data to simulate the behavior of human perception, it is expected to be applicable to a wide range of accents for which speech data is scarce.

语音识别口音鲁棒自监督离散表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。