构建118小时多语言低语数据集,解决真实录音难问题
WhispSynth: Scaling Multilingual Whisper Corpus through Real Data Curation and A Novel Pitch-free Generative Framework
- 用无音高生成框架+语音合成技术,从真实数据中重建高保真低语
- 产出479名说话人共118小时高质量低语数据,音色与语义保留完整
- 适合语音合成、低语识别研究者使用,支持高自然度语音生成
低语语音生成受限于数据收集困难,因低语声波振幅小,高保真录制极难实现。本文提出WhispSynth,一种基于新型高保真生成框架的大规模多语言低语语料库。具体采用融合基于可微信号处理(DDSP)的无音高方法与文本到语音(TTS)模型的流水线,将包括新构建的WhispNJU数据集在内的多种资源整合,生成了来自479名说话人的118小时高保真低语语音。与传统合成或含噪真实数据不同,该数据引擎忠实保留源声带音色与语言内容,同时保证声学一致性,为文本到低语研究提供坚实基础。实验表明,WhispSynth质量显著优于现有语料库;使用WhispSynth训练的CosyWhisper模型,在语音自然度上达到与真实样本相当水平。官方代码与资源已公开于https://github.com/tan90xx/cosywhisper。
原文摘要 · Abstract (English)
Whisper generation is constrained by the difficulty of data collection. Because whispered speech has low acoustic amplitude, high-fidelity recording is challenging. In this paper, we introduce WhispSynth, a large-scale multilingual corpus constructed via a novel high-fidelity generative framework. Specifically, we propose a pipeline integrating Differentiable Digital Signal Processing (DDSP)-based pitch-free method with Text-to-Speech (TTS) models. This framework refines a comprehensive collection of resources, including our newly constructed WhispNJU dataset, into 118 hours of high-fidelity whispered speech from 479 speakers. Unlike standard synthetic or noisy real data, our data engine faithfully preserves source vocal timbre and linguistic content while ensuring acoustic consistency, providing a robust foundation for text-to-whisper research. Experimental results demonstrate that WhispSynth exhibits significantly higher quality than existing corpora. Moreover, our CosyWhisper, tuned with WhispSynth, achieves speech naturalness on par with ground-truth samples. The official implementation and related resources are available at https://github.com/tan90xx/cosywhisper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。