分析SSL特征如何识别语音伪造模型,揭示其优缺点。
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
- 通过控制生成参数,研究SSL特征对模型架构的识别能力。
- 发现微小变化如提示词、声码器会影响识别准确率。
- 适合关注语音伪造溯源安全的研究者和开发者。
语音伪造模型溯源旨在通过识别生成给定音频样本的源头模型,防止合成语音被滥用,实现责任追溯并为厂商提供信息支持。该任务极具挑战性,但自监督学习(SSL)提取的声学特征已展现出最先进的溯源能力,其成功背后的驱动因素及判别能力的边界仍不明确。本文系统研究了SSL特征如何捕捉语音伪造中的架构指纹。通过控制音频生成过程的多个维度,我们揭示了模型检查点、文本提示、声码器或说话人身份的细微变化如何影响溯源效果。研究结果为基于SSL的语音伪造溯源在真实场景下的鲁棒性、偏差与局限性提供了新见解,凸显了其优势与脆弱性。
原文摘要 · Abstract (English)
Audio deepfake model attribution aims to mitigate the misuse of synthetic speech by identifying the source model responsible for generating a given audio sample, enabling accountability and informing vendors. The task is challenging, but self-supervised learning (SSL)-derived acoustic features have demonstrated state-of-the-art attribution capabilities, yet the underlying factors driving their success and the limits of their discriminative power remain unclear. In this paper, we systematically investigate how SSL-derived features capture architectural signatures in audio deepfakes. By controlling multiple dimensions of the audio generation process we reveal how subtle perturbations in model checkpoints, text prompts, vocoders, or speaker identity influence attribution. Our results provide new insights into the robustness, biases, and limitations of SSL-based deepfake attribution, highlighting both its strengths and vulnerabilities in realistic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。