arXiv:2501.17615cs.CLcs.SD2025-01被引 1

用跨语言嵌入聚类提升低资源多语种语音识别准确率

Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition

  • 通过跨语言嵌入聚类构建分层软最大值解码器
  • 在15个语言的简化数据集上显著提升低资源语言识别准确率
  • 适合多语种语音识别与资源匮乏语言研究者

我们提出一种聚焦自动语音识别(ASR)解码阶段的新方法,以增强多语种性能,尤其针对低资源语言。该方法利用跨语言嵌入聚类构建分层软最大值(H-Softmax)解码器,使不同语言中相似的词元共享相近的解码表示。该方法克服了以往基于霍夫曼树的H-Softmax方法依赖浅层特征进行词元相似性评估的局限。在包含15种语言的下采样数据集上的实验表明,该方法能有效提升低资源多语种语音识别的准确性。

原文摘要 · Abstract (English)

We present a novel approach centered on the decoding stage of Automatic Speech Recognition (ASR) that enhances multilingual performance, especially for low-resource languages. It utilizes a cross-lingual embedding clustering method to construct a hierarchical Softmax (H-Softmax) decoder, which enables similar tokens across different languages to share similar decoder representations. It addresses the limitations of the previous Huffman-based H-Softmax method, which relied on shallow features in token similarity assessments. Through experiments on a downsampled dataset of 15 languages, we demonstrate the effectiveness of our approach in improving low-resource multilingual ASR accuracy.

语音识别多语言低资源嵌入聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。