arXiv:2501.00328cs.SDcs.CL2025-01中稿 · 2025 IEEE Internat…被引 4

构建首个大规模越南语多风格语音识别数据集,解决跨风格识别难题

VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition

  • 基于公开数据自动构建18.7万条语音,覆盖1406人多风格语料
  • 单风格训练模型在跨风格测试中性能下降超20%,加入新数据后显著提升
  • 适合研究越南语语音识别、跨风格泛化及低资源语言建模者

近期语音识别研究聚焦于注册与测试语音间差异带来的脆弱性问题,尤其是多风格现象下不同语音风格导致的挑战。现有越南语语音识别资源或规模有限,或缺乏风格多样性,致使多风格影响的研究尚未展开。本文提出VoxVietnam,首个面向越南语语音识别的大规模多风格数据集,包含超过18.7万条来自1406名说话人的语音,并设计自动化流程从公开来源大规模构建该数据集。实验表明,仅在单一风格上训练的模型在跨风格测试中面临显著性能下降;而将VoxVietnam纳入训练后,模型表现大幅提升。本研究系统评估了多风格现象带来的挑战,并验证了使用该数据集进行多风格训练的有效性。

原文摘要 · Abstract (English)

Recent research in speaker recognition aims to address vulnerabilities due to variations between enrolment and test utterances, particularly in the multi-genre phenomenon where the utterances are in different speech genres. Previous resources for Vietnamese speaker recognition are either limited in size or do not focus on genre diversity, leaving studies in multi-genre effects unexplored. This paper introduces VoxVietnam, the first multi-genre dataset for Vietnamese speaker recognition with over 187,000 utterances from 1,406 speakers and an automated pipeline to construct a dataset on a large scale from public sources. Our experiments show the challenges posed by the multi-genre phenomenon to models trained on a single-genre dataset, and demonstrate a significant increase in performance upon incorporating the VoxVietnam into the training process. Our experiments are conducted to study the challenges of the multi-genre phenomenon in speaker recognition and the performance gain when the proposed dataset is used for multi-genre training.

语音识别多风格越南语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。