分析语音识别中说话人间与说话人内差异的权衡关系,指导数据收集。
Study on Inter and Intra Speaker Variability in Speaker Recognition
- 基于VoxTube数据集,研究训练数据中说话人间与说话人内差异的关系
- 揭示说话人数量与会话多样性之间的优化平衡点
- 提供媒体平台数据采集建议,适合语音识别研究者参考
在现代神经网络语音识别系统中,优化说话人数量与时间变异性的权衡关系至关重要,同时需确保数据收集在时间上可行。本文基于VoxTube数据集,针对无文本依赖的语音识别任务,分析了训练数据中说话人间与说话人内变异性的依赖关系。此外,本工作还首次公开了VoxTube数据集中每条语音片段的上传日期元数据,旨在为从媒体托管平台收集和筛选数据提供指南与最佳实践,助力研究人员更高效地构建语音识别系统。
原文摘要 · Abstract (English)
Optimization of a trade-off between the number of speakers and their temporal variability (or session diversity) is crucial for the development of a speaker recognition system together with making the data collection process feasible from a time perspective. In this article, we provide the analysis of dependency between inter and intra speaker variability in training data for the modern neural network-based speaker recognition system using the VoxTube dataset for text-independent speaker recognition task. Besides, an auxiliary contribution of this work is a release of upload date metadata per utterance in a VoxTube dataset. We want this article to contribute to guidelines and best practices for collecting and filtering data from media hosting platforms to facilitate the efforts of researchers in developing speaker recognition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。