区分真假歌声伪造,先筛低质再识歌手,提升识别准确率。
Not All Deepfakes Are Created Equal: Triaging Audio Forgeries for Robust Deepfake Singer Identification
- 分两阶段:先用判别器筛除低质伪造,再用真实录音训练模型识别
- 在真实与高质伪造音频上均优于现有方法,识别准确率显著提升
- 适合音乐版权保护、艺人形象维权等场景使用
高度逼真的歌唱声线深度伪造的泛滥对保护艺术家形象和内容真实性构成了严峻挑战。在语音深度伪造中自动识别歌手是一种有前景的手段,可帮助艺术家和权利持有者防范其声音被未经授权使用,但该问题仍处于开放研究阶段。基于‘最具危害的伪造通常质量最高’这一前提,我们提出一个两阶段流程来识别歌手的声线特征。首先通过判别模型筛选出无法准确再现声线特征的低质量伪造。随后,一个仅在真实录音上训练的模型,用于识别剩余高质量伪造及真实音频中的歌手。实验表明,该系统在真实与合成内容上均持续优于现有基线方法。
原文摘要 · Abstract (English)
The proliferation of highly realistic singing voice deepfakes presents a significant challenge to protecting artist likeness and content authenticity. Automatic singer identification in vocal deepfakes is a promising avenue for artists and rights holders to defend against unauthorized use of their voice, but remains an open research problem. Based on the premise that the most harmful deepfakes are those of the highest quality, we introduce a two-stage pipeline to identify a singer's vocal likeness. It first employs a discriminator model to filter out low-quality forgeries that fail to accurately reproduce vocal likeness. A subsequent model, trained exclusively on authentic recordings, identifies the singer in the remaining high-quality deepfakes and authentic audio. Experiments show that this system consistently outperforms existing baselines on both authentic and synthetic content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。