arXiv:2504.06753cs.SDcs.AI2025-04AAAI被引 20

提出新型音频伪造检测方法,可识别语音、歌声、音乐等全类型伪造音频。

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

论文配图:Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception
图 1 · 摘自论文原文
  • 通过小参数量提示调优提升自监督学习模型对伪造音频的感知能力。
  • 在跨类型检测任务中平均错误率仅3.58%,显著优于传统微调方法。
  • 适合需要高泛化性音频安全防护的场景,如社交媒体审核与智能语音系统。

音频生成技术的快速发展带来了语音、声音、歌唱声和音乐等多类型深度伪造音频的威胁,严重破坏多媒体安全与可信度。现有反制措施在单一类型检测中表现良好,但在跨类型场景下性能下降。本文首次全面建立全类型深度伪造音频检测(All-Type ADD)基准,涵盖语音、声音、歌唱声和音乐四类伪造音频的跨类型检测任务。提出提示调优自监督学习(PT-SSL)训练范式,通过学习专用提示令牌优化自监督前端,参数量仅为微调(FT)的1/458。针对不同音频类型的听觉感知特性,设计小波提示调优(WPT-SSL)方法,在不增加训练参数的前提下,从频域捕捉跨类型伪造特征,提升全类型检测性能。通过联合训练所有类型伪造音频构建通用反制模型。实验表明,WPT-XLSR-AASIST在所有测试集上平均等错误率(EER)达3.58%,表现最优。

原文摘要 · Abstract (English)

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the all-type ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL front-end by learning specialized prompt tokens for ADD, requiring 458x fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types, we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets.

音频伪造深度学习自监督检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。