arXiv:2409.14743eess.AScs.SD2024-09被引 27

用大模型生成真假混杂语音,测试反伪造系统的实际抗攻击能力

LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation

  • 用大模型和克隆技术生成130小时混合真假语音数据
  • 现有检测系统在未见场景下最差误判率高达24.49%
  • 揭示系统对特定语音模型的偏见,适合安全研究者参考

以往伪造语音数据集多从防御方视角构建,用于开发反伪造系统,却未考虑攻击者的多样动机。为更贴近真实场景,我们构建了LlamaPartialSpoof——一个130小时的语音数据集,包含完全伪造与部分伪造语音,利用大语言模型(LLM)和语音克隆技术评估反伪造系统鲁棒性。通过分析攻击者与防御者的关键信息,我们发现当前反伪造系统存在若干可被利用的漏洞,如对特定文本转语音模型或拼接方法的偏倚。实验表明,现有伪造语音检测系统在未见过的场景中泛化能力差,最佳等错误率(EER)仅为24.49%。

原文摘要 · Abstract (English)

Previous fake speech datasets were constructed from a defender's perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially fake speech, using a large language model (LLM) and voice cloning technologies to evaluate the robustness of CMs. By examining valuable information for both attackers and defenders, we identify several key vulnerabilities in current CM systems, which can be exploited to enhance attack success rates, including biases toward certain text-to-speech models or concatenation methods. Our experimental results indicate that the current fake speech detection system struggle to generalize to unseen scenarios, achieving a best performance of 24.49% equal error rate.

语音伪造对抗攻击大模型应用数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。