用零样本语音合成生成数据,提升个性化语音增强效果
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
- 通过零样本语音合成生成个性化语音数据
- 合成数据显著提升个性化语音增强模型性能
- 适合语音增强与生成交叉研究者参与
本文提出一项新挑战,要求在ICASSP 2025生成式数据增强研讨会中,使用零样本文本到语音(TTS)系统为下游个性化语音增强(PSE)任务生成语音数据。由于隐私顾虑及测试场景录音的技术难度,获取高质量个性化数据十分困难。为此,利用生成模型合成数据受到广泛关注。本挑战要求参赛者首先构建零样本TTS系统以扩充个性化数据,随后用该增强数据训练PSE系统。通过此挑战,我们旨在探究零样本TTS生成数据的质量对PSE模型性能的影响。我们还提供了基于开源零样本TTS模型的基线实验,以鼓励参与并建立基准。相关代码与模型检查点已公开。
原文摘要 · Abstract (English)
This paper presents a new challenge that calls for zero-shot text-to-speech (TTS) systems to augment speech data for the downstream task, personalized speech enhancement (PSE), as part of the Generative Data Augmentation workshop at ICASSP 2025. Collecting high-quality personalized data is challenging due to privacy concerns and technical difficulties in recording audio from the test scene. To address these issues, synthetic data generation using generative models has gained significant attention. In this challenge, participants are tasked first with building zero-shot TTS systems to augment personalized data. Subsequently, PSE systems are asked to be trained with this augmented personalized dataset. Through this challenge, we aim to investigate how the quality of augmented data generated by zero-shot TTS models affects PSE model performance. We also provide baseline experiments using open-source zero-shot TTS models to encourage participation and benchmark advancements. Our baseline code implementation and checkpoints are available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。