首个大规模粤语语音数据集,支持多维标注。
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
- 构建六模块流水线,实现多维度粤语语音标注
- 覆盖21,800小时跨10个领域,含声学质量与说话人属性
- 适配ASR/TTS训练,助力粤语语音技术发展
语音理解与生成的发展得益于大规模高质量语音数据集的出现。其中,自动语音识别(ASR)和文本转语音(TTS)是最成熟的基础任务。然而,全球约8490万母语者使用的粤语,因标注资源匮乏,导致其ASR与TTS性能不佳。为此,我们提出WenetSpeech-Pipe,一套面向语音理解与生成的大型语音语料库构建流水线,包含音频采集、说话人属性标注、语音质量标注、自动语音识别、文本后处理及识别结果投票六大模块,实现丰富高质量的多维标注。基于此,我们发布WenetSpeech-Yue,首个具有多维标注的大规模粤语语音语料库,涵盖21,800小时、10个领域,包含ASR转录、文本置信度、说话人身份、年龄、性别、语音质量评分等标注。同时发布WSYue-eval,包含两个组件:用于短/长句、代码切换、复杂声学环境评估的WSYue-ASR-eval,以及用于标准与泛化测试的WSYue-TTS-eval(含基础与覆盖子集)。实验表明,在WenetSpeech-Yue上训练的模型可达到媲美现有最优(SOTA)粤语ASR与TTS系统(包括商业与大语言模型基系统)的性能,凸显本数据集与流水线的价值。
原文摘要 · Abstract (English)
The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. It comprises six modules: Audio Collection, Speaker Attributes Annotation, Speech Quality Annotation, Automatic Speech Recognition, Text Postprocessing and Recognizer Output Voting, enabling rich and high-quality annotations. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。