构建跨合成器与说话人迁移的音频伪造检测基准,评估模型真实场景泛化能力。
TwinShift: Benchmarking Audio Deepfake Detection across Synthesizer and Speaker Shifts
- 设计六种独立合成系统+互斥说话人集,模拟真实世界未知生成与说话人变化
- 实验揭示现有检测器在跨模型和跨说话人时性能显著下降,平均准确率降低超20%
- 为开发更鲁棒的音频伪造检测系统提供可复现的评估标准和改进方向
音频深度伪造威胁日益严重,已被用于诈骗和虚假信息传播。当前主要挑战在于检测器能否抵御未见过的生成方法和多样化说话人,因生成技术迭代迅速。尽管现有系统在基准测试中表现良好,但其在新条件下的泛化能力仍不足,限制了实际应用可靠性。为此,我们提出 TWINSHIFT 基准,专为评估检测器在严格未知条件下的鲁棒性而设计。该基准基于六种不同的合成系统,每种系统均配以互不重叠的说话人集合,从而实现对生成模型与说话人身份同时变化时泛化能力的严格评估。通过大量实验,TWINSHIFT 揭示了重要的鲁棒性差距,暴露出被忽视的局限性,并为构建更可靠的音频深度伪造检测(ADD)系统提供系统性指导。TWINSHIFT 基准可通过 https://github.com/intheMeantime/TWINSHIFT 获取。
原文摘要 · Abstract (English)
Audio deepfakes pose a growing threat, already exploited in fraud and misinformation. A key challenge is ensuring detectors remain robust to unseen synthesis methods and diverse speakers, since generation techniques evolve quickly. Despite strong benchmark results, current systems struggle to generalize to new conditions limiting real-world reliability. To address this, we introduce TWINSHIFT, a benchmark explicitly designed to evaluate detection robustness under strictly unseen conditions. Our benchmark is constructed from six different synthesis systems, each paired with disjoint sets of speakers, allowing for a rigorous assessment of how well detectors generalize when both the generative model and the speaker identity change. Through extensive experiments, we show that TWINSHIFT reveals important robustness gaps, uncover overlooked limitations, and provide principled guidance for developing ADD systems. The TWINSHIFT benchmark can be accessed at https://github.com/intheMeantime/TWINSHIFT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。