arXiv:2506.19014cs.SDcs.AI2025-06被引 2

首个聚焦南亚英语口音的语音深度伪造检测数据集,填补了跨文化识别空白。

IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection

  • 构建包含50名印度英语母语者的真实与伪造音频,覆盖27.17小时
  • 在南亚语境下检测性能超越ASVspoof21(DF),挑战性高于ITW基准
  • 公开可获取,适合语音安全、跨文化语音识别研究者使用

语音深度伪造技术虽带来助手机器人、无障碍沟通等益处,但严重威胁数字通信安全与信任。现有数据集缺乏多样族裔口音,难以应对南亚等多语言文化场景。本文提出IndieFake Dataset(IFD),包含50名印度英语说话者的27.17小时真实与伪造音频,具备均衡分布与说话人级标注,弥补ASVspoof21(DF)等数据集不足。在IFD上评估多种基线模型,结果表明其性能优于ASVspoof21(DF),且比标准In-The-Wild(ITW)数据集更具挑战性。完整数据集及参考片段已公开供研究使用。

原文摘要 · Abstract (English)

Advancements in audio deepfake technology offers benefits like AI assistants, better accessibility for speech impairments, and enhanced entertainment. However, it also poses significant risks to security, privacy, and trust in digital communications. Detecting and mitigating these threats requires comprehensive datasets. Existing datasets lack diverse ethnic accents, making them inadequate for many real-world scenarios. Consequently, models trained on these datasets struggle to detect audio deepfakes in diverse linguistic and cultural contexts such as in South-Asian countries. Ironically, there is a stark lack of South-Asian speaker samples in the existing datasets despite constituting a quarter of the worlds population. This work introduces the IndieFake Dataset (IFD), featuring 27.17 hours of bonafide and deepfake audio from 50 English speaking Indian speakers. IFD offers balanced data distribution and includes speaker-level characterization, absent in datasets like ASVspoof21 (DF). We evaluated various baselines on IFD against existing ASVspoof21 (DF) and In-The-Wild (ITW) datasets. IFD outperforms ASVspoof21 (DF) and proves to be more challenging compared to benchmark ITW dataset. The complete dataset, along with documentation and sample reference clips, is publicly accessible for research use on project website.

语音伪造数据集安全检测南亚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。