对比主动水印与被动检测,统一评估标准
A Comparative Study on Proactive and Passive Detection of Deepfake Speech
- 构建统一框架,公平比较主动水印与被动检测方法
- 在相同数据集上测试,发现不同模型对语音畸变敏感度不同
- 适合安全研究者和防御系统设计者参考
针对深度伪造语音的防御方案分为两类:主动水印模型与被动检测器。两者虽应对相同威胁,但训练、优化和评估方式差异导致难以统一评价。本文提出一个统一框架,对两类模型进行深度伪造语音检测评估。为确保公平性,所有模型均在相同数据集上训练与测试,并采用统一指标评估性能。同时分析其对各类对抗攻击的鲁棒性,发现不同模型对不同语音属性畸变存在差异化脆弱性。训练与评估代码已开源。
原文摘要 · Abstract (English)
Solutions for defending against deepfake speech fall into two categories: proactive watermarking models and passive conventional deepfake detectors. While both address common threats, their differences in training, optimization, and evaluation prevent a unified protocol for joint evaluation and selecting the best solutions for different cases. This work proposes a framework to evaluate both model types in deepfake speech detection. To ensure fair comparison and minimize discrepancies, all models were trained and tested on common datasets, with performance evaluated using a shared metric. We also analyze their robustness against various adversarial attacks, showing that different models exhibit distinct vulnerabilities to different speech attribute distortions. Our training and evaluation code is available at Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。