通过增强语音伪造中的伪影,显著提升反欺骗检测效果
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
- 用噪声注入+语音增强,放大伪造语音中的隐藏伪影
- 在ASVspoof2019上提升检测性能44.44%,2021年数据集上提升26.34%
- 不依赖特定模型,适配多种语音增强和反欺骗架构
伪造语音总是包含生成模型引入的伪影。尽管已有多种反欺骗方法提出,但多数集中在架构改进上。本文研究伪造语音中伪影为何隐蔽,并探索如何增强其存在感。提出一种模型无关的流水线:通过噪声添加、噪声提取与伪影放大三步,强化伪造信号中的异常特征。该方法可兼容不同语音增强模型与反欺骗架构。实验表明,在ASVspoof2019上检测性能提升最高达44.44%,在ASVspoof2021上提升26.34%。
原文摘要 · Abstract (English)
Spoofed utterances always contain artifacts introduced by generative models. While several countermeasures have been proposed to detect spoofed utterances, most primarily focus on architectural improvements. In this work, we investigate how artifacts remain hidden in spoofed speech and how to enhance their presence. We propose a model-agnostic pipeline that amplifies artifacts using speech enhancement and various types of noise. Our approach consists of three key steps: noise addition, noise extraction, and noise amplification. First, we introduce noise into the raw speech. Then, we apply speech enhancement to extract the entangled noise and artifacts. Finally, we amplify these extracted features. Moreover, our pipeline is compatible with different speech enhancement models and countermeasure architectures. Our method improves spoof detection performance by up to 44.44\% on ASVspoof2019 and 26.34\% on ASVspoof2021.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。