构建首个覆盖多类伪造语音的综合数据集并提出高效检测模型
Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
- 设计跨算法合成语音的统一数据集,涵盖真实、合成与部分伪造样本
- 检测准确率83.55%(mAP),片段级误报率仅1.07%且F1达92.19%
- 单模型实现真伪判断、位置定位与生成算法识别,无需后处理
由于虚假信息和身份冒用的风险日益加剧,辨别合成语音与真实语音变得至关重要。尽管已有多种合成语音数据集被开发,但它们通常聚焦于特定领域,限制了全面研究的应用。为填补这一空白,我们提出了Speech-Forensics数据集,广泛覆盖真实、合成及部分伪造语音样本,包含由多种高质量算法生成的多个片段。同时,我们提出一种名为TEST的时序语音定位网络,旨在无需复杂后处理的情况下,同时完成真实性检测、多段伪造区域定位及合成算法识别。TEST通过融合LSTM与Transformer提取更强的时序语音表征,并在多尺度金字塔特征上采用密集预测以估计合成区间。模型在语句级别达到83.55%的平均mAP和5.25%的EER;在片段级别,其EER为1.07%,F1得分为92.19%。这些结果凸显了模型在合成语音综合分析中的强大能力,为该领域的未来研究与实际应用提供了有力支持。
原文摘要 · Abstract (English)
Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas, limiting their utility for comprehensive research. To fill this gap, we propose the Speech-Forensics dataset by extensively covering authentic, synthetic, and partially forged speech samples that include multiple segments synthesized by different high-quality algorithms. Moreover, we propose a TEmporal Speech LocalizaTion network, called TEST, aiming at simultaneously performing authenticity detection, multiple fake segments localization, and synthesis algorithms recognition, without any complex post-processing. TEST effectively integrates LSTM and Transformer to extract more powerful temporal speech representations and utilizes dense prediction on multi-scale pyramid features to estimate the synthetic spans. Our model achieves an average mAP of 83.55% and an EER of 5.25% at the utterance level. At the segment level, it attains an EER of 1.07% and a 92.19% F1 score. These results highlight the model's robust capability for a comprehensive analysis of synthetic speech, offering a promising avenue for future research and practical applications in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。