用DiffWave模型生成高保真婴儿哭声,提升声音多样性。
Towards the Synthesis of Non-speech Vocalizations
- 基于DiffWave框架实现无条件婴儿哭声生成
- 在Baby Chillanto和deBarbaro数据集上生成高质量新哭声
- 适合音频合成与人机交互研究者参考
本报告聚焦于使用DiffWave框架进行婴儿哭声的无条件生成,该框架在从噪声中生成高质量音频方面表现优异。采用两个独立的婴儿哭声数据集——Baby Chillanto和deBarbaro——训练DiffWave模型,以生成具有高保真度和多样性的新哭声。研究重点在于评估DiffWave在无条件生成任务中的能力,验证其在非语音发声合成中的适用性与有效性。
原文摘要 · Abstract (English)
In this report, we focus on the unconditional generation of infant cry sounds using the DiffWave framework, which has shown great promise in generating high-quality audio from noise. We use two distinct datasets of infant cries: the Baby Chillanto and the deBarbaro cry dataset. These datasets are used to train the DiffWave model to generate new cry sounds that maintain high fidelity and diversity. The focus here is on DiffWave's capability to handle the unconditional generation task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。