通过噪声增强录音文件,提升远场语音识别性能
oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
- 在注册音频中添加噪声,对齐测试与注册数据源
- ResNet模型达0.75的DCF和12.79%的EER
- 适用于远场语音识别场景,尤其适合机器人语音系统
本文提出一种新型数据增强方法,通过在注册音频中加入噪声,有效对齐测试与注册音频的声源特性,提升两者可比性。实验采用多种预训练模型,其中ResNet表现最佳,原始性能为0.84的DCF和13.44%的EER;经该增强技术后,性能提升至0.75 DCF和12.79% EER。对比分析显示,ResNet优于ECPA、Mel-spectrogram、Payonnet及Titanet large等模型。该方法及其不同增强策略显著推动了RoboVox远场说话人识别系统的成功。
原文摘要 · Abstract (English)
In this study, we address the challenge of speaker recognition using a novel data augmentation technique of adding noise to enrollment files. This technique efficiently aligns the sources of test and enrollment files, enhancing comparability. Various pre-trained models were employed, with the resnet model achieving the highest DCF of 0.84 and an EER of 13.44. The augmentation technique notably improved these results to 0.75 DCF and 12.79 EER for the resnet model. Comparative analysis revealed the superiority of resnet over models such as ECPA, Mel-spectrogram, Payonnet, and Titanet large. Results, along with different augmentation schemes, contribute to the success of RoboVox far-field speaker recognition in this paper
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。