arXiv:2601.19767cs.SD2026-01中稿 · ICASSP 2026被引 2

用可微分k均值建模跨语言语音可懂度优势,提升带口音语音识别准确率。

Advanced Modeling of Interlanguage Speech Intelligibility Benefit with L1-L2 Multi-Task Learning Using Differentiable K-Means for Accent-Robust Discrete Token-Based ASR

  • 通过可微分k均值联合优化母语与第二语言语音识别任务
  • 在有限口音数据下实现约20%的识别准确率相对提升
  • 适合需要鲁棒口音识别的多语言语音系统研发者

构建对外国口音语音具有鲁棒性的自动语音识别(ASR)系统是当今全球化世界的重要挑战。已有研究通过复现跨语言语音可懂度优势(ISIB)现象——即非母语说话者口音对同母语听众比对母语听众更易理解——来提升基于音素标记的ASR在口音语音上的表现。该现象通过使用说话人母语(L1)在自监督学习(SSL)特征空间中学习k均值聚类中心以获取音素标记实现。本研究提出一种更先进的ISIB建模方法:采用可微分k均值,并联合优化整个模块在L1和L2 ASR任务上的性能。实验表明,该方法在仅使用母语语音时优于基线,在额外引入少量口音语音时更是取得约20%的相对识别准确率提升。

原文摘要 · Abstract (English)

Building ASR systems robust to foreign-accented speech is an important challenge in today's globalized world. A prior study explored the way to enhance the performance of phonetic token-based ASR on accented speech by reproducing the phenomenon known as interlanguage speech intelligibility benefit (ISIB), where foreign-accented speech is more intelligible to listeners sharing the speaker's native language than to native listeners. ISIB was technically implemented by using the speaker's L1 to learn k-means cluster centroids in an SSL feature space to obtain phonetic tokens. In this study, we propose a more advanced modeling of ISIB. By employing differentiable k-means and optimizing the entire module for both L1 and L2 ASR, the proposed method outperformed the baselines, both when using only native speech and when additionally incorporating a limited amount of accented speech. Notably, in the latter scenario, our method achieved approximately a 20% relative improvement in recognition accuracy.

语音识别口音鲁棒多任务学习可微分聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。