首个真实场景的演唱语音增强基准,推动该领域评估与进步
SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
- 构建首个涵盖多样真实场景的演唱语音增强数据集
- 发现感知质量与可懂度存在持续权衡关系
- 在真实演唱数据上训练能显著提升性能且不损害语音能力
本文提出一个演唱语音增强的基准。当前该领域发展受限于缺乏真实可用的评估数据。为填补这一空白,本文首次引入SingVERSE,首个面向真实世界场景的演唱语音增强基准,覆盖多种声学环境,并提供成对的高质量录音参考。基于SingVERSE,我们对前沿模型进行了全面评估,发现感知质量与可懂度之间存在一致的权衡。最后,研究显示在域内演唱数据上训练可显著提升增强效果,且不降低语音处理能力,为该领域提供了简单而有效的改进路径。本工作为社区提供了一个基础性基准及关键洞见,以推动这一被忽视领域的未来发展。
原文摘要 · Abstract (English)
This paper presents a benchmark for singing voice enhancement. The development of singing voice enhancement is limited by the lack of realistic evaluation data. To address this gap, this paper introduces SingVERSE, the first real-world benchmark for singing voice enhancement, covering diverse acoustic scenarios and providing paired, studio-quality clean references. Leveraging SingVERSE, we conduct a comprehensive evaluation of state-of-the-art models and uncover a consistent trade-off between perceptual quality and intelligibility. Finally, we show that training on in-domain singing data substantially improves enhancement performance without degrading speech capabilities, establishing a simple yet effective path forward. This work offers the community a foundational benchmark together with critical insights to guide future advances in this underexplored domain. Demopage: https://singverse.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。