用语音预训练模型识别音频伪造源的语调特征,提升溯源精度。
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
- 采用多种先进语音预训练模型捕捉语音语调特征,用于伪造音频溯源。
- x-vector模型虽参数最少,却在基准数据集上表现最优。
- 提出FINDER融合策略,结合Whisper与x-vector实现当前最佳性能。
本文研究多种前沿语音预训练模型(PTMs)在捕捉音频深度伪造源语调特征方面的能力,以提升音频深度伪造溯源(ADSD)性能。语调特征是各生成源的独特标识,模型捕捉能力越强,溯源效果越好。实验在ASVSpoof 2019和CFAD两个基准数据集上进行,评估了多个在不同语调任务中表现优异的SOTA PTMs。结果表明,尽管参数量最小,用于说话人识别的x-vector模型表现最佳,可能因其预训练任务更擅长捕捉源的语调特性。受语音识别与伪造检测中多模型融合提升性能的启发,本文提出FINDER方法,实现对Whisper与x-vector表示的有效融合。该融合方案在所有单个模型及基线融合方法中表现最优,达到当前最先进水平。
原文摘要 · Abstract (English)
In this work, we investigate various state-of-the-art (SOTA) speech pre-trained models (PTMs) for their capability to capture prosodic signatures of the generative sources for audio deepfake source attribution (ADSD). These prosodic characteristics can be considered one of major signatures for ADSD, which is unique to each source. So better is the PTM at capturing prosodic signs better the ADSD performance. We consider various SOTA PTMs that have shown top performance in different prosodic tasks for our experiments on benchmark datasets, ASVSpoof 2019 and CFAD. x-vector (speaker recognition PTM) attains the highest performance in comparison to all the PTMs considered despite consisting lowest model parameters. This higher performance can be due to its speaker recognition pre-training that enables it for capturing unique prosodic characteristics of the sources in a better way. Further, motivated from tasks such as audio deepfake detection and speech recognition, where fusion of PTMs representations lead to improved performance, we explore the same and propose FINDER for effective fusion of such representations. With fusion of Whisper and x-vector representations through FINDER, we achieved the topmost performance in comparison to all the individual PTMs as well as baseline fusion techniques and attaining SOTA performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。