arXiv:2602.00443cs.SDcs.MM2026-02被引 2

评测语音克隆在真实场景下的鲁棒性,发现多个系统性弱点。

RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models

  • 构建覆盖多语言、长文本等场景的综合评测集
  • 18个模型测试显示内容一致性与抗干扰能力普遍不足
  • 适合研究语音合成安全与鲁棒性的学者使用

现代语音克隆(零样本文本到语音,TTS)可仅用数秒参考音频生成接近目标说话人的语音,广泛应用于个性化语音接口和配音。但在实际部署中,常面临噪声参考音频、不完整文本提示、多语言与长文本生成、后处理及对抗扰动等问题,影响系统鲁棒性。尽管编解码器令牌语言模型和基于扩散的TTS快速发展,真实场景下的鲁棒性仍缺乏系统评估。本文提出RVCBench,一个全面的语音克隆鲁棒性评测数据集与基准。RVCBench涵盖受控文本-音频匹配、多语言与长文本、表达性提示、后处理条件以及被动或主动音频扰动的任务对齐测试。在18项鲁棒性评估、225名说话人、14,370条语音上,支持对输入敏感度、生成稳定性、输出韧性、扰动鲁棒性、说话人相似度及深度伪造检测性的统一评估。我们评测了18个代表性开源语音克隆模型,揭示内容一致性、说话人相似度、长文本稳定性、后处理韧性、对抗鲁棒性及检测器可区分性等方面的系统性缺陷。代码与数据集已公开,以支持可复现评估及未来研究。代码:https://github.com/Nanboy-Ronan/RVCBench。数据集:https://huggingface.co/datasets/Nanboy/RVCBench。

原文摘要 · Abstract (English)

Modern voice cloning, also known as zero-shot text-to-speech (TTS), can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing. In practice, these systems often face noisy reference audio, imperfect text prompts, multilingual and long-form generation, post-processing, and adversarial perturbations, all of which can weaken robustness. Despite rapid progress in codec-token language models and diffusion-based TTS, robustness under realistic deployment shifts remains underexplored. This paper introduces RVCBench, a comprehensive dataset and benchmark for evaluating robustness in voice cloning. RVCBench provides task-aligned tests covering controlled text-audio pairing, multilingual and long-form scenarios, expressive prompts, post-processing conditions, and passive or proactive audio perturbations. Across 18 robustness evaluations, 225 speakers, and 14,370 utterances, RVCBench supports unified evaluation of input sensitivity, generation stability, output resilience, perturbation robustness, speaker similarity, and deepfake detectability. We evaluate 18 representative open-source voice cloning models and reveal systematic vulnerabilities in content consistency, speaker similarity, long-form stability, post-processing resilience, adversarial robustness, and detector-facing separability. We release the code and dataset to support reproducible evaluation and future research on robust voice cloning, speech synthesis, and audio generation. Code: https://github.com/Nanboy-Ronan/RVCBench. Dataset: https://huggingface.co/datasets/Nanboy/RVCBench.

语音克隆鲁棒性评测深度伪造TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。