针对音频部分伪造问题,构建新数据集并提出分组件检测方法。
CompSpoof: A Dataset and Joint Learning Framework for Component-Level Audio Anti-spoofing Countermeasures
- 分离语音与环境音,分别进行反伪造检测。
- 在多组合伪造场景下,检测准确率显著提升。
- 适合需要高精度音频真伪识别的安防与语音系统。
组件级音频伪造(Comp-Spoof)是一种新型音频篡改技术,仅对信号中的特定成分(如语音或环境声)进行伪造,而其他成分保持真实。现有反伪造数据集和方法将整个语段视为完全真实或完全伪造,无法有效检测此类局部篡改。为此,本文构建了新数据集CompSpoof,涵盖语音与环境声的多种真实与伪造组合。同时提出一种增强分离的联合学习框架:先分离音频组件,再对每个组件独立应用反伪造模型,并通过联合学习保留关键检测信息。大量实验表明,该方法优于基线,验证了分组件检测的必要性与有效性。数据集与代码已公开于https://github.com/XuepingZhang/CompSpoof。
原文摘要 · Abstract (English)
Component-level audio Spoofing (Comp-Spoof) targets a new form of audio manipulation where only specific components of a signal, such as speech or environmental sound, are forged or substituted while other components remain genuine. Existing anti-spoofing datasets and methods treat an utterance or a segment as entirely bona fide or entirely spoofed, and thus cannot accurately detect component-level spoofing. To address this, we construct a new dataset, CompSpoof, covering multiple combinations of bona fide and spoofed speech and environmental sound. We further propose a separation-enhanced joint learning framework that separates audio components apart and applies anti-spoofing models to each one. Joint learning is employed, preserving information relevant for detection. Extensive experiments demonstrate that our method outperforms the baseline, highlighting the necessity of separate components and the importance of detecting spoofing for each component separately. Datasets and code are available at: https://github.com/XuepingZhang/CompSpoof.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。