构建首个孟加拉语深度伪造语音数据集,助力低资源语言反伪造研究。
BanglaFake: Constructing and Evaluating a Specialized Bengali Deepfake Audio Dataset
- 用顶尖语音合成模型生成高保真伪造语音,覆盖真实与伪造语料
- 12,260条真实语音与13,260条伪造语音,经母语者评估自然度得分3.40
- 为孟加拉语深度伪造检测提供关键数据支撑,填补低资源语言空白
由于缺乏数据集和细微的声学特征,低资源语言如孟加拉语的深度伪造语音检测面临挑战。为此,我们提出BanglaFake,一个包含12,260条真实语音和13,260条深度伪造语音的孟加拉语深度伪造音频数据集。合成语音采用前沿文本转语音(TTS)模型生成,确保高自然度与质量。通过定性与定量分析评估该数据集:30名母语者参与的平均意见分(MOS)显示,鲁棒自然度得分为3.40,可懂度得分为4.01。基于梅尔频率倒谱系数(MFCCs)的t-SNE可视化揭示了真实与伪造语音区分困难的问题。该数据集为推进孟加拉语深度伪造检测提供了关键资源,弥补了低资源语言研究的不足。
原文摘要 · Abstract (English)
Deepfake audio detection is challenging for low-resource languages like Bengali due to limited datasets and subtle acoustic features. To address this, we introduce BangalFake, a Bengali Deepfake Audio Dataset with 12,260 real and 13,260 deepfake utterances. Synthetic speech is generated using SOTA Text-to-Speech (TTS) models, ensuring high naturalness and quality. We evaluate the dataset through both qualitative and quantitative analyses. Mean Opinion Score (MOS) from 30 native speakers shows Robust-MOS of 3.40 (naturalness) and 4.01 (intelligibility). t-SNE visualization of MFCCs highlights real vs. fake differentiation challenges. This dataset serves as a crucial resource for advancing deepfake detection in Bengali, addressing the limitations of low-resource language research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。