用非语义音频特征提升伪造语音检测的泛化能力。
Generalizable Audio Spoofing Detection using Non-Semantic Representations
- 利用非语义通用音频表示捕捉底层声学特征
- 在域外数据上显著优于现有方法,尤其在公开数据集表现突出
- 适合需要跨场景鲁棒检测的应用场景
生成模型的快速发展使合成语音制作变得容易,导致基于语音的服务面临伪造攻击风险。现有深度伪造检测方法常因泛化能力不足,在真实数据上表现急剧下降。本文提出一种基于非语义通用音频表示的可泛化伪造检测方法。通过TRILL和TRILLsson模型筛选合适的非语义特征,实验表明该方法在域内测试集上性能相当,而在域外测试集上显著优于当前最优方法。尤其在公开数据集上表现优异,超越了基于手工特征、语义嵌入及端到端架构的方法。
原文摘要 · Abstract (English)
Rapid advancements in generative modeling have made synthetic audio generation easy, making speech-based services vulnerable to spoofing attacks. Consequently, there is a dire need for robust countermeasures more than ever. Existing solutions for deepfake detection are often criticized for lacking generalizability and fail drastically when applied to real-world data. This study proposes a novel method for generalizable spoofing detection leveraging non-semantic universal audio representations. Extensive experiments have been performed to find suitable non-semantic features using TRILL and TRILLsson models. The results indicate that the proposed method achieves comparable performance on the in-domain test set while significantly outperforming state-of-the-art approaches on out-of-domain test sets. Notably, it demonstrates superior generalization on public-domain data, surpassing methods based on hand-crafted features, semantic embeddings, and end-to-end architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。