通过分层注意力与对比学习,提升语音伪造检测的跨域泛化能力。
HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
- 构建分层注意力机制,捕捉帧、层及层组间的依赖关系
- 在ASVspoof 2021 DF和In-the-Wild上分别达到1.93%和6.87%的EER
- 适用于检测跨生成方式与录制条件的语音伪造
由现代文本转语音(TTS)与语音转换系统生成的语音伪造正变得越来越难以与真实语音区分,严重威胁安全与网络信任。尽管最先进的自监督模型提供了丰富的多层表征,现有检测器仍独立处理各层,忽视了对识别合成伪影至关重要的时序与层级依赖。我们提出HierCon,一种结合基于边距的对比学习的分层注意力框架,可建模时间帧、相邻层及层组间的依赖关系,同时促进领域不变嵌入。在ASVspoof 2021 DF与In-the-Wild数据集上的评估表明,该方法达到领先性能(1.93%与6.87% EER),相比独立层加权分别提升36.6%与22.5%。结果与注意力可视化证实,分层建模增强了对跨领域生成技术与录音条件的泛化能力。
原文摘要 · Abstract (English)
Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich multi-layer representations, existing detectors treat layers independently and overlook temporal and hierarchical dependencies critical for identifying synthetic artefacts. We propose HierCon, a hierarchical layer attention framework combined with margin-based contrastive learning that models dependencies across temporal frames, neighbouring layers, and layer groups, while encouraging domain-invariant embeddings. Evaluated on ASVspoof 2021 DF and In-the-Wild datasets, our method achieves state-of-the-art performance (1.93% and 6.87% EER), improving over independent layer weighting by 36.6% and 22.5% respectively. The results and attention visualisations confirm that hierarchical modelling enhances generalisation to cross-domain generation techniques and recording conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。