分析原始波形模型在语音识别中的错误模式,发现其混淆规律与传统方法一致。
Phonetic Error Analysis of Raw Waveform Acoustic Models

- 基于SincNet与双向LSTM的模型实现13.9%/15.3%电话错误率
- 迁移学习使辅音错误率下降约三倍,元音改善较小
- 错误模式反映语音固有相似性,适用于语音系统优化
我们对TIMIT数据集上的原始波形声学模型进行了超出整体电话错误率(PER)的错误模式分析。通过三种广义音素类别(BPC)划分,构建替换错误的混淆矩阵。所用模型结合参数化(SincNet、Sinc2Net)或非参数化CNN与双向LSTM,分别在开发集/测试集上达到13.9%/15.3%的PER,为当前原始波形模型在TIMIT上的最佳结果。从WSJ迁移学习后PER降至11.3%/12.3%,优于滤波器组基线。按BPC分析显示,双向LSTM对过渡依赖类别的提升最显著,而迁移学习对辅音的改善约为元音的三倍。混淆模式在原始波形与滤波器组系统间高度一致,表明主要混淆源于语音本身的相似性。
原文摘要 · Abstract (English)
We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER). PER is decomposed across three broad phonetic class (BPC) categorisations, and confusion matrices are constructed from substitution errors. Our models combine parametric (SincNet, Sinc2Net) or non-parametric CNNs with Bidirectional LSTMs, achieving 13.9%/15.3% PER on Dev/Test, the best reported results for raw waveform models on TIMIT. Transfer learning from WSJ reduces PER to 11.3%/12.3%, surpassing the Filterbank baseline. Per-BPC analysis reveals that BLSTM layers benefit transition-dependent classes most, while WSJ transfer learning improves consonants roughly three times more than vowels. Confusion patterns are consistent across raw waveform and Filterbank systems, indicating that the dominant confusions reflect inherent phonetic similarities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。