提出SASTNet,提升编码器生成假语音的溯源泛化能力
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
- 融合语义与声学特征,用Whisper和Wav2vec2+AudioMAE联合编码
- 在CodecFake+数据集上达到当前最优溯源准确率
- 解决仅用重编码数据训练导致的过拟合问题,适合真实场景应用
针对基于编码器的深度伪造语音(CodecFake)的源溯源任务,现有方法性能不佳。本文指出,仅用编码器重编码数据训练的模型易在非语音区域过拟合,难以泛化至真实生成音频。为此,提出语义-声学溯源网络SASTNet,联合使用Whisper进行语义特征提取,以及Wav2vec2与AudioMAE进行声学特征编码。在CodecFake+数据集的CoSG测试集上,SASTNet实现当前最佳性能,验证了其在真实场景下可靠溯源的有效性。
原文摘要 · Abstract (English)
Recent attempts at source tracing for codec-based deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-the-art performance on the CoSG test set of the CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。