用多比例重建残差融合提升跨域语音伪造检测效果
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

- 用冻结的DiT模型在不同遮蔽率下生成残差图
- 在ASVspoof5和ITW数据集上分别达6.54%和13.84%的错误等价率
- 适合关注跨域迁移与音频伪造检测的研究者
语音伪造检测器在生成器、语料或录制条件变化时性能常下降。本文使用仅在真实语音上训练的扩散变压器(DiT)作为冻结的重建探测器,在遮蔽率0.5、0.75和0.9下生成多比例残差图。由于残差具有领域敏感性,所提出的音频锚定检测器将冻结的WavLM声学表示直接投影至融合求和路径,不进行门控衰减,仅将残差作为标量门控加性修正。预设种子42的实验在ASVspoof 5 Eval上取得6.5442% EER / 0.18456 min-DCF,ITW Full上为13.8372% / 0.36921;三种子均值分别为6.8885% (0.3308) 和15.3328% (2.0719%)。后者在两种监督设置下均优于独立优化的WavLM-ResNet18参考模型。辅助监督使动态竞争融合的平均ITW EER从18.4007%提升至25.2968%,反而恶化了三个种子结果。结果支持重建残差作为互补证据,并推动一种非竞争性听觉路径用于从ASVspoof 5到ITW的迁移,但未声称锚定机制本身的因果贡献。
原文摘要 · Abstract (English)
Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. Reconstructions at masking ratios 0.5, 0.75, and 0.9 yield explicit multi-ratio residual maps. Because these residuals are domain sensitive, our audio-anchored detector passes the projected frozen-WavLM auditory representation into the fusion sum without gate-based attenuation and uses residuals only as a scalar-gated additive correction. The pre-specified seed-42 run obtains 6.5442% EER / 0.18456 min-DCF on ASVspoof 5 Eval and 13.8372% / 0.36921 on ITW Full; three-seed means are 6.8885 (0.3308)% and 15.3328 (2.0719)%. The latter is below a separately optimized WavLM-ResNet18 reference under both supervision settings. Auxiliary supervision raises dynamic competitive fusion from 18.4007% to 25.2968% mean ITW EER, worsening all three seeds. The results support reconstruction residuals as complementary evidence and motivate a non-competitive auditory path for ASVspoof 5-to-ITW transfer, without claiming a componentwise causal ablation of anchoring alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。