用耳蜗图特征提升嘈杂环境下的语音识别准确率
Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
- 用128通道的gammatone滤波器生成耳蜗图特征
- 在-5dB信噪比下训练模型,识别准确率显著提升
- 适合语音识别在噪声、混响和失真环境中的应用
说话人识别旨在从已知说话人集合中通过声音识别个体。环境噪声、混响和失真会导致提取特征退化,降低自动说话人识别(SID)系统的性能。本文提出一种在噪声、不匹配、混响及失真环境下均鲁棒的噪声自适应说话人识别系统。该方法利用耳蜗图(cochleagram)提取说话人特征,采用128通道gammatone滤波器组(频率范围50–8000 Hz)生成二维耳蜗图。使用宽带与窄带噪声叠加清洁语音,在不同信噪比(SNR)条件下生成含噪耳蜗图。仅将-5 dB SNR的清洁与含噪耳蜗图输入卷积神经网络(CNN),构建噪声自适应说话人模型(NASM)。该模型在特定噪声下训练后,可在清洁及多种噪声环境下评估。此外,还测试了系统在混响和失真数据上的鲁棒性。结果表明,该系统相比现有基于神经图的SID系统,识别准确率有明显提升。
原文摘要 · Abstract (English)
Speaker Identification refers to the process of identifying a person using one's voice from a collection of known speakers. Environmental noise, reverberation and distortion make the task of automatic speaker identification challenging as extracted features get degraded thus affecting the performance of the speaker identification (SID) system. This paper proposes a robust noise adapted SID system under noisy, mismatched, reverberated and distorted environments. This method utilizes an auditory features called cochleagram to extract speaker characteristics and thus identify the speaker. A $128$ channel gammatone filterbank with a frequency range from $50$ to $8000$ Hz was used to generate 2-D cochleagrams. Wideband as well as narrowband noises were used along with clean speech to obtain noisy cochleagrams at various levels of signal to noise ratio (SNR). Both clean and noisy cochleagrams of only $-5$ dB SNR were then fed into a convolutional neural network (CNN) to build a speaker model in order to perform SID which is referred as noise adapted speaker model (NASM). The NASM was trained using a certain noise and then was evaluated using clean and various types of noises. Moreover, the robustness of the proposed system was tested using reverberated as well as distorted test data. Performance of the proposed system showed a measurable accuracy improvement over existing neurogram based SID system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。