多骨干自监督集成模型实现高精度语音伪造检测,揭示生成与检测间的不对称性。
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry
- 采用四种语音模型自监督集成,提升伪造语音检测能力。
- 检测准确率达100%,生成语音在对抗检测中仍能逃过超半数识别系统。
- 揭示生成与检测能力不匹配问题,适合安全与反伪造研究者参考。
本文介绍团队「Go-To-Germany」在ImageCLEF 2026语音伪造检测与生成任务中的参与情况。我们的检测系统基于四骨干自监督学习(SSL)集成,融合WavLM-Large、Wav2Vec2-XLS-R-300M、ECAPA-TDNN和x-vector表示,在官方评估中获得0.9522的最终得分,对参赛者生成的伪造语音达到1.0000的完美准确率,对主办方未公开的真实数据集得分为0.8875。在生成子任务中,我们提交的F5-TTS v1基线经统一混响处理,作为刻意反取证探测,以0.4304的最终得分(词错误率WER 4.99%,字符错误率CER 2.07%)排名第一。论文详述了四个生成模型(GLM-TTS、F5-TTS、XTTS v2、CosyVoice3)的程序,其中官方提交版本由此衍生。我们进行跨任务分析,揭示显著不对称:检测系统可识别100%参赛者生成的伪造语音,而我们的生成模型虽在生成任务中排名第一,却仅能逃避61.4%和56.2%的参赛者与主办方检测器。通过留一speaker交叉验证(LOSO)、三区域骨干结构、置信区间与主成分分析等伪造性消融实验,支持多骨干集成的架构保险假说。此外,提供五项跨任务洞察及五项预注册伪造实验,连接生成侧逃避行为与检测侧设计决策,并公开报告11.25%的假阳性差距为部署主要挑战。
原文摘要 · Abstract (English)
This paper describes the participation of team "Go-To-Germany" in the ImageCLEF 2026 Audio Deepfake Detection and Generation task. Our detection system, built on a four-backbone self-supervised learning (SSL) ensemble combining WavLM-Large, Wav2Vec2-XLS-R-300M, ECAPA-TDNN, and x-vector representations, achieved a final score of 0.9522 on the official ImageCLEF 2026 evaluation, with perfect accuracy (1.0000) on participant-generated deepfakes and 0.8875 on the held-out organizer ground-truth real data. For the Generation sub-task, our official team submission, an F5-TTS v1 baseline processed with a uniform reverberation pass and submitted as a deliberate anti-forensic probe, ranked first with a final score of 0.4304 (word error rate (WER) 4.99%, character error rate (CER) 2.07%); details of our four-model program (GLM-TTS, F5-TTS, XTTS v2, CosyVoice3), from which the official entry was drawn, appear in the paper. We present a cross-track analysis revealing a pronounced asymmetry: our detection system identifies 100% of participant-generated deepfakes, while our official generation entry, despite ranking first in the Audio Generation sub-task and evading 61.4% and 56.2% of participant and organizer detectors, attains a Final Score of 0.4304 against 0.9522 on the Detection side. We further report falsification-based ablation experiments (LOSO 56-speaker cross-validation, three-region backbone geometry, bootstrap confidence intervals, and PCA analysis) that motivate our architectural-insurance hypothesis for multi-backbone SSL ensembling. We complement these results with five cross-track insights and five pre-registered falsification experiments connecting generation-side evasion to detection-side design decisions, and we openly report an 11.25% false-positive gap on held-out organizer real recordings as the principal open challenge for deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。