提升音频伪造系统溯源在开放场景下的准确率,效果显著。
Open-Set Source Tracing of Audio Deepfake Systems
- 提出新方法Softmax Energy(SME)改进异常检测
- 实现FPR95降至8.3%,比原有方法提升31%
- 适合对抗新型未知伪造系统的安全研究者
现有音频伪造系统溯源研究多聚焦封闭集场景,而针对开放集性能评估的样本极少。随着新型音频伪造系统不断涌现,具备鲁棒性的开放集溯源能力至关重要。本文采用Interspeech 2025源追溯专项赛的评测协议,提出一种用于分布外(OOD)检测的能量分数新变体——软最大能量(Softmax Energy, SME)。实验表明,将传统温度缩放能量分数替换为SME,可使标准FPR95(真阳性率为95%时的假阳性率)指标平均相对提升31%。进一步结合SME引导训练及复制合成、编解码器和混响增强策略,最终达到FPR95为8.3%的性能。该方法在识别未知伪造系统方面表现优异。
原文摘要 · Abstract (English)
Existing research on source tracing of audio deepfake systems has focused primarily on the closed-set scenario, while studies that evaluate open-set performance are limited to a small number of unseen systems. Due to the large number of emerging audio deepfake systems, robust open-set source tracing is critical. We leverage the protocol of the Interspeech 2025 special session on source tracing to evaluate methods for improving open-set source tracing performance. We introduce a novel adaptation to the energy score for out-of-distribution (OOD) detection, softmax energy (SME). We find that replacing the typical temperature-scaled energy score with SME provides a relative average improvement of 31% in the standard FPR95 (false positive rate at true positive rate of 95%) measure. We further explore SME-guided training as well as copy synthesis, codec, and reverberation augmentations, yielding an FPR95 of 8.3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。