自动提取文本发音关联,提升语音识别上下文准确性。
Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing
- 基于语音-文本对齐生成发音关联,无需人工词典
- 在普通话数据上显著提升端到端语音识别性能
- 适合无标准发音词典的方言或语言场景
在语言声学中,有效区分不同书面文本间的发音关联是一个重要问题。传统方法依赖人工设计的发音词典获取发音关联。本文提出一种数据驱动的自动文本发音关联(ATPC)方法。该方法所需监督信息与端到端语音识别(E2E-ASR)系统训练一致,即语音与对应文本标注。首先使用迭代训练的时间戳估计器(ITSE)算法对齐语音与其对应的文本符号;然后通过语音编码器将语音转换为语音嵌入;最后比较不同文本符号对应的语音嵌入距离,获得ATPC。在普通话上的实验结果表明,ATPC能有效提升E2E-ASR在上下文偏置中的性能,且在缺乏人工发音词典的方言或语言中具有应用前景。
原文摘要 · Abstract (English)
Effectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually designed pronunciation lexicons. In this paper, we propose a data-driven method to automatically acquire these pronunciation correlations, called automatic text pronunciation correlation (ATPC). The supervision required for this method is consistent with the supervision needed for training end-to-end automatic speech recognition (E2E-ASR) systems, i.e., speech and corresponding text annotations. First, the iteratively-trained timestamp estimator (ITSE) algorithm is employed to align the speech with their corresponding annotated text symbols. Then, a speech encoder is used to convert the speech into speech embeddings. Finally, we compare the speech embeddings distances of different text symbols to obtain ATPC. Experimental results on Mandarin show that ATPC enhances E2E-ASR performance in contextual biasing and holds promise for dialects or languages lacking artificial pronunciation lexicons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。