首个面向吴语的8000小时语音数据集与评估基准,推动方言语音技术发展
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
- 构建8000小时多维度标注的吴语语音数据集
- 推出覆盖6大任务的标准化评估基准Wu-Bench
- 开源多任务模型,助力方言语音智能研究
低资源方言语音处理仍是构建包容性、鲁棒语音技术的根本挑战。尽管吴语汉语具有重要语言学意义和庞大的使用者群体,长期受限于缺乏大规模语音数据、标准化评估基准及公开可用模型。本文提出WenetSpeech-Wu,首个大规模、多维度标注的开源吴语语音语料库,包含约8000小时多样化的语音数据。基于此数据集,我们引入WenetSpeech-Wu-Bench,首个标准化且公开可访问的吴语语音处理评估基准,涵盖自动语音识别(ASR)、吴语到普通话翻译、说话人属性预测、语音情感识别、文本转语音(TTS)合成及指令跟随式语音合成(instruct TTS)六大任务。此外,我们发布了一系列在WenetSpeech-Wu上训练的强健开源模型,在多个任务中实现竞争力表现,实证验证了所提数据集的有效性。这些贡献共同奠定了全面的吴语语音处理生态系统基础,所有数据集、基准和模型均开源,以支持未来方言语音智能研究。
原文摘要 · Abstract (English)
Speech processing for low-resource dialects remains a fundamental challenge in developing inclusive and robust speech technologies. Despite its linguistic significance and large speaker population, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models. In this work, we present WenetSpeech-Wu, the first large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect, comprising approximately 8,000 hours of diverse speech data. Building upon this dataset, we introduce WenetSpeech-Wu-Bench, the first standardized and publicly accessible benchmark for systematic evaluation of Wu dialect speech processing, covering automatic speech recognition (ASR), Wu-to-Mandarin translation, speaker attribute prediction, speech emotion recognition, text-to-speech (TTS) synthesis, and instruction-following TTS (instruct TTS). Furthermore, we release a suite of strong open-source models trained on WenetSpeech-Wu, establishing competitive performance across multiple tasks and empirically validating the effectiveness of the proposed dataset. Together, these contributions lay the foundation for a comprehensive Wu dialect speech processing ecosystem, and we open-source proposed datasets, benchmarks, and models to support future research on dialectal speech intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。