分析2025年歌声转换挑战赛结果,揭示风格迁移仍难实现。
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
- 构建新数据集与双任务框架,评估33个系统
- 顶级系统保留歌手身份,但风格自然度仍不足
- 客观指标相关性不够,无法替代人工评分
本文对最新一届歌声转换挑战赛(Singing Voice Conversion Challenge 2025)的成果进行深入分析。该赛事旨在对比和理解不同语音转换系统在受控环境下的表现。与以往仅关注歌手身份转换不同,本届新增了对演唱风格转换的关注。为确保评估严谨,研究团队构建了新数据库,设立两项任务,开源基线模型,并开展为期两个月的大规模众包听觉测试与客观评价。共评估了33个系统。听觉测试结果显示,顶级系统在歌手身份保真度上接近真实样本。然而,在建模呼吸感、滑音、颤音等动态演唱风格方面仍面临挑战,导致自然度不足。进一步分析指出,传统相似性测试与动态偏好测试在评估风格相似性方面存在局限。计算皮尔逊等级相关系数表明,依赖性客观指标(如音高对齐)和非匹配指标(如说话人嵌入)与主观评分相关性最高,但仍不足以完全替代主观评价。
原文摘要 · Abstract (English)
We present a thorough analysis of the findings of the latest iteration of the Singing Voice Conversion Challenge, a scientific event aiming to compare and understand different voice conversion systems in a controlled environment. Compared to previous iterations which solely focused on converting the singer identity, this year we also focused on converting the singing style of the singer. To create a controlled environment and thorough evaluations, we developed a new challenge database, introduced two tasks, open-sourced baselines, and conducted large-scale crowd-sourced listening tests and objective evaluations. The challenge was run for two months and in total we evaluated 33 different systems. The results of the large-scale crowd-sourced listening test showed that top systems had comparable singer identity scores to ground truth samples. However, modeling the singing style and consequently achieving high naturalness still remains a challenge in this task, primarily due to the difficulty in modeling dynamic information in breathy, glissando, and vibrato singing styles. Further analyses of the challenge also discuss the limitations of both the traditional similarity test and the dynamic preference test in evaluating singing style similarity. Moreover, calculating Spearman's rank correlation coefficient shows that dependent objective metrics such as chroma-alignment and non-match metrics such as speaker embeddings are the most correlated to subjective scores, but are still not at a level where it could be considered as a true replacement for subjective scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。