用小波包与梅尔倒谱结合,提升无文本语音识别鲁棒性
Speaker Recognition -- Wavelet Packet Based Multiresolution Feature Extraction Approach
- 融合梅尔倒谱与小波包多分辨率特性提取声纹特征
- 在VoxForge和TIMIT数据集上识别与验证准确率均优于传统方法
- 在不同信噪比噪声下仍保持稳定性能,适合真实环境应用
本文提出一种基于小波包的新型无文本语音识别特征提取方法,通过结合梅尔频率倒谱系数(MFCC)与小波包变换(WPT)实现。该方法利用MFCC模拟人耳听觉特性,同时融合WPT的多分辨率分析能力与抗噪优势。采用高斯混合模型(GMM)和隐马尔可夫模型(HMM)分别作为说话人识别与验证的分类器。在VoxForge语音语料库和CSTR US KED Timit数据库上进行测试,并在不同信噪比(SNR)条件下加入标准噪声以评估抗噪性能。实验结果表明,该方法在说话人识别与验证任务中均取得更优效果。
原文摘要 · Abstract (English)
This paper proposes a novel Wavelet Packet based feature extraction approach for the task of text independent speaker recognition. The features are extracted by using the combination of Mel Frequency Cepstral Coefficient (MFCC) and Wavelet Packet Transform (WPT).Hybrid Features technique uses the advantage of human ear simulation offered by MFCC combining it with multi-resolution property and noise robustness of WPT. To check the validity of the proposed approach for the text independent speaker identification and verification we have used the Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM) respectively as the classifiers. The proposed paradigm is tested on voxforge speech corpus and CSTR US KED Timit database. The paradigm is also evaluated after adding standard noise signal at different level of SNRs for evaluating the noise robustness. Experimental results show that better results are achieved for the tasks of both speaker identification as well as speaker verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。