通过层选择与融合提升音频伪造检测在未知场景下的泛化能力
Improving Out-of-Domain Audio Deepfake Detection via Layer Selection and Fusion of SSL-Based Countermeasures
- 分析不同预训练模型各层表现,自动选取最优层进行特征提取
- 单层选择使参数减少80%且性能接近多层融合,同时提升对未知数据的适应性
- 多编码器得分级融合显著增强对未知伪造攻击的鲁棒性,适合实际部署
基于冻结预训练自监督学习(SSL)编码器的音频伪造检测系统,在结合层加权池化方法(如多头因子化注意力池化,MHFA)时表现出较高性能,但仍难以泛化至域外(OOD)条件。本文研究六种不同预训练SSL模型在四个测试语料上的表现,进行逐层分析以确定贡献最大的层。对比单层策略与自动层选择(如MHFA)发现,优选最佳层可获得优异结果,同时将系统参数减少高达80%。不同测试语料和编码器间性能差异显著,表明编码器预训练策略影响显著。最终,多个编码器在得分层面的融合有效提升了对域外攻击的泛化能力。
原文摘要 · Abstract (English)
Audio deepfake detection systems based on frozen pre-trained self-supervised learning (SSL) encoders show a high level of performance when combined with layer-weighted pooling methods, such as multi-head factorized attentive pooling (MHFA). However, they still struggle to generalize to out-of-domain (OOD) conditions. We tackle this problem by studying the behavior of six different pre-trained SSLs, on four different test corpora. We perform a layer-by-layer analysis to determine which layers contribute most. Next, we study the pooling head, comparing a strategy based on a single layer with automatic selection via MHFA. We observed that selecting the best layer gave very good results, while reducing system parameters by up to 80%. A wide variation in performance as a function of test corpus and SSL model is also observed, showing that the pre-training strategy of the encoder plays a role. Finally, score-level fusion of several encoders improved generalization to OOD attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。