arXiv:2409.05924cs.SDeess.AS2024-09被引 12

用Transformer检测音频伪造,能持续学习新造假手法。

Continuous Learning of Transformer-based Audio Deepfake Detection

  • 基于音频频谱变换器构建检测模型,提升准确率。
  • 在多个基准数据集上表现优异,少量标注数据即可更新模型。
  • 适合需要实时应对新型语音伪造的安防与内容审核场景。

本文提出一种新型音频深度伪造检测框架,旨在实现两个目标:一是对已有伪造音频达到最高检测准确率;二是以少样本学习方式有效进行持续学习。我们使用多种深度音频生成方法大规模收集伪造音频数据,并通过压缩、远场录音、噪声等增强手段提升数据多样性。采用音频频谱变换器(Audio Spectrogram Transformer)作为检测模型,在多个基准数据集上取得良好性能。此外,设计了一个持续学习插件模块,仅需极少标注数据即可高效更新模型,显著优于传统微调方法。

原文摘要 · Abstract (English)

This paper proposes a novel framework for audio deepfake detection with two main objectives: i) attaining the highest possible accuracy on available fake data, and ii) effectively performing continuous learning on new fake data in a few-shot learning manner. Specifically, we conduct a large audio deepfake collection using various deep audio generation methods. The data is further enhanced with additional augmentation methods to increase variations amidst compressions, far-field recordings, noise, and other distortions. We then adopt the Audio Spectrogram Transformer for the audio deepfake detection model. Accordingly, the proposed method achieves promising performance on various benchmark datasets. Furthermore, we present a continuous learning plugin module to update the trained model most effectively with the fewest possible labeled data points of the new fake type. The proposed method outperforms the conventional direct fine-tuning approach with much fewer labeled data points.

音频伪造Transformer持续学习少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。