仅用音频和名人名,无需时间标注就能训练出顶尖语音嵌入模型。
State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
- 只用音频和名人姓名,通过弱监督方式训练说话人嵌入模型。
- 在语音验证任务上达到与有监督训练相当的顶尖性能。
- 可处理未知说话人片段,适合大规模弱标签数据训练。
本文改进并验证了一种基于弱标注训练说话人嵌入提取器的方法。具体而言,仅使用源VoxCeleb视频的音频流和名人姓名,而不依赖其在录音中的出现时间区间。我们实验了基于ResNet和WavLM的超参数与嵌入提取器,结果表明该方法在说话人验证任务中达到与在VoxCeleb数据集上标准有监督训练相当的性能。此外,我们扩展了方法,将与名人共现的未知说话人片段也纳入训练,通常这些片段会被丢弃。该方法无需说话人时间戳和多模态对齐,使得大规模弱标签语音数据可直接用于训练顶尖嵌入模型,为VoxCeleb类数据集构建提供无需视觉信息的替代方案。
原文摘要 · Abstract (English)
In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments belonging to unknown speakers appearing alongside the celebrities, which are typically discarded. Removing the need for speaker timestamps and multimodal alignment, our method unlocks the use of large-scale weakly labeled speech data, enabling direct training of state-of-the-art embedding extractors and offering a visual-free alternative to VoxCeleb-style dataset creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。