arXiv:2509.22808cs.CL2025-09被引 2

首个多方言阿拉伯语语音伪造检测数据集,助力对抗合成语音攻击

ArFake: A Robust Framework for Multi-Dialect Arabic Speech Spoofing Detection Benchmark

  • 构建首个多方言阿拉伯语伪造语音数据集,融合多种生成模型样本
  • FishSpeech在卡萨布兰卡语料库上生成的语音最逼真,难辨真假
  • 结合人评分与语音识别错误率,科学评估合成语音欺骗性

随着文本转语音生成模型的兴起,区分真实与合成语音变得愈发困难,尤其针对研究较少的阿拉伯语及其众多方言。现有伪造检测工作主要集中在英语,对阿拉伯语存在显著空白。本文提出首个多方言阿拉伯语伪造语音数据集。为评估各生成模型的合成音频难度并筛选最具挑战性的样本,我们设计了包含两类方法的评估流程:基于现代嵌入的分类器结合分类头,以及基于MFCC特征的传统机器学习算法;同时采用RawNet2架构进行对比。流程还引入基于人工评分的平均意见分(MOS),并通过自动语音识别模型计算词错误率(WER)来衡量合成语音的可理解性。结果表明,FishSpeech在卡萨布兰卡语料库上的阿拉伯语语音克隆表现最佳,生成的合成语音更真实、更具欺骗性。但仅依赖单一生成模型可能影响数据集的泛化能力。

原文摘要 · Abstract (English)

With the rise of generative text-to-speech models, distinguishing between real and synthetic speech has become challenging, especially for Arabic that have received limited research attention. Most spoof detection efforts have focused on English, leaving a significant gap for Arabic and its many dialects. In this work, we introduce the first multi-dialect Arabic spoofed speech dataset. To evaluate the difficulty of the synthesized audio from each model and determine which produces the most challenging samples, we aimed to guide the construction of our final dataset either by merging audios from multiple models or by selecting the best-performing model, we conducted an evaluation pipeline that included training classifiers using two approaches: modern embedding-based methods combined with classifier heads; classical machine learning algorithms applied to MFCC features; and the RawNet2 architecture. The pipeline further incorporated the calculation of Mean Opinion Score based on human ratings, as well as processing both original and synthesized datasets through an Automatic Speech Recognition model to measure the Word Error Rate. Our results demonstrate that FishSpeech outperforms other TTS models in Arabic voice cloning on the Casablanca corpus, producing more realistic and challenging synthetic speech samples. However, relying on a single TTS for dataset creation may limit generalizability.

语音伪造阿拉伯语TTS检测多方言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。