arXiv:2512.21702cs.SDcs.AI2025-12中稿 · publication in 202…被引 2

针对孟加拉语深度伪造音频,提出首个系统性检测方案。

Zero-Shot to Zero-Lies: Detecting Bengali Deepfake Audio through Transfer Learning

  • 用预训练模型零样本检测,效果有限
  • 微调后ResNet18准确率达79.17%,性能显著提升
  • 为低资源语言语音伪造检测提供有效范式

语音合成与语音转换技术的快速发展使深度伪造音频成为重大安全威胁。孟加拉语深度伪造检测研究仍处于空白。本文基于BanglaFake数据集,评估多个预训练模型在孟加拉语深度伪造音频上的零样本检测能力,包括Wav2Vec2-XLSR-53、Whisper、PANNsCNN14、WavLM和Audio Spectrogram Transformer。零样本结果表现不佳,最佳模型Wav2Vec2-XLSR-53准确率为53.80%,AUC为56.60%,EER为46.20%。随后对多种架构进行微调,包括Wav2Vec2-Base、LCNN、LCNN-Attention、ResNet18、ViT-B16和CNN-BiLSTM。微调后模型性能大幅提升,其中ResNet18达到最高准确率79.17%、F1分数79.12%、AUC 84.37%、EER 24.35%。实验验证微调显著优于零样本推理。本研究建立了首个孟加拉语深度伪造音频检测的系统性基准,证明了微调深度学习模型在低资源语言中的有效性。

原文摘要 · Abstract (English)

The rapid growth of speech synthesis and voice conversion systems has made deepfake audio a major security concern. Bengali deepfake detection remains largely unexplored. In this work, we study automatic detection of Bengali audio deepfakes using the BanglaFake dataset. We evaluate zeroshot inference with several pretrained models. These include Wav2Vec2-XLSR-53, Whisper, PANNsCNN14, WavLM and Audio Spectrogram Transformer. Zero-shot results show limited detection ability. The best model, Wav2Vec2-XLSR-53, achieves 53.80% accuracy, 56.60% AUC and 46.20% EER. We then f ine-tune multiple architectures for Bengali deepfake detection. These include Wav2Vec2-Base, LCNN, LCNN-Attention, ResNet18, ViT-B16 and CNN-BiLSTM. Fine-tuned models show strong performance gains. ResNet18 achieves the highest accuracy of 79.17%, F1 score of 79.12%, AUC of 84.37% and EER of 24.35%. Experimental results confirm that fine-tuning significantly improves performance over zero-shot inference. This study provides the first systematic benchmark of Bengali deepfake audio detection. It highlights the effectiveness of f ine-tuned deep learning models for this low-resource language.

语音伪造孟加拉语微调安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。