评测零样本对话语音克隆,推动自然对话语音生成技术发展
The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings
- 设计双赛道:开放与受限数据使用,评估不同条件下的语音克隆能力
- 发布100小时高质量对话语音数据集,支持研究与模型训练
- 揭示当前零样本语音克隆在自然语调和对话连贯性上的关键挑战
ISCSLP 2024 对话语音克隆(CoVoC)挑战旨在基准测试并推进零样本自发风格语音克隆技术,尤其关注对话语音中的自然行为生成。挑战包含两个赛道:无限制赛道允许自由使用数据与模型,受限赛道仅限使用公开开源数据集。本次挑战同步发布了一个100小时的高质量对话语音数据集。本文详述了数据集、赛道设置、参赛系统、评估结果及主要发现。
原文摘要 · Abstract (English)
The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge aims to benchmark and advance zero-shot spontaneous style voice cloning, particularly focusing on generating spontaneous behaviors in conversational speech. The challenge comprises two tracks: an unconstrained track without limitation on data and model usage, and a constrained track only allowing the use of constrained open-source datasets. A 100-hour high-quality conversational speech dataset is also made available with the challenge. This paper details the data, tracks, submitted systems, evaluation results, and findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。