arXiv:2510.19414eess.AScs.AI2025-10被引 2

构建120小时真实场景语音伪造数据集,提升防伪检测实战能力。

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection

  • 融合零样本语音合成与真实设备回放录音,模拟实际攻击环境。
  • 现有模型在回放音频上准确率降至59.6%,暴露出泛化缺陷。
  • 新数据集使检测模型平均错误率更低,适合部署于真实场景。

语音深度伪造日益泛滥,尤其在电话诈骗和身份盗用等真实场景中引发严重担忧。尽管许多反欺骗系统在实验室生成的合成语音上表现良好,但在面对物理回放攻击——一种常见且低成本的实际攻击方式时往往失效。实验表明,基于现有数据集训练的模型在回放音频上的平均准确率降至59.6%。为此,我们提出EchoFake,一个涵盖超过13,000名说话人、总计120小时以上的音频数据集,包含前沿的零样本文本转语音(TTS)语音及在多种设备和真实环境条件下采集的物理回放录音。此外,我们评估了三种基线检测模型,结果表明在EchoFake上训练的模型在多个数据集上实现更低的平均等错误率(EER),展现出更强的泛化能力。通过引入更贴近真实部署的挑战,EchoFake为推进语音伪造检测技术提供了更真实的基准。

原文摘要 · Abstract (English)

The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing systems have demonstrated promising performance on lab-generated synthetic speech, they often fail when confronted with physical replay attacks-a common and low-cost form of attack used in practical settings. Our experiments show that models trained on existing datasets exhibit severe performance degradation, with average accuracy dropping to 59.6% when evaluated on replayed audio. To bridge this gap, we present EchoFake, a comprehensive dataset comprising more than 120 hours of audio from over 13,000 speakers, featuring both cutting-edge zero-shot text-to-speech (TTS) speech and physical replay recordings collected under varied devices and real-world environmental settings. Additionally, we evaluate three baseline detection models and show that models trained on EchoFake achieve lower average EERs across datasets, indicating better generalization. By introducing more practical challenges relevant to real-world deployment, EchoFake offers a more realistic foundation for advancing spoofing detection methods.

语音伪造数据集反欺骗真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。