构建大规模音效数据集,提升深度伪造音频检测能力
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
- 整合7种文本转音频模型生成音效,构建4.3万段音视频样本
- 包含2.6万段合成音效与1.7万段真实音效,支持细粒度检测研究
- 专为孤立音效深度伪造检测设计,适合对抗性测试与评估
尽管音频深度伪造检测技术已取得显著进展,但现有代表性检测器在面对合成音效时泛化能力有限。现有环境音频数据集如EnvSDD虽具初步价值,但在规模和生成来源上仍不足以支撑对孤立音效深度伪造的系统研究。为此,本文提出SynSFX,一个包含43,374段音频片段(其中26,452段为合成,16,922段为真实)的大规模数据集,覆盖7种主流文本转音频模型,旨在推动针对合成音效深度伪造的检测与评估研究。
原文摘要 · Abstract (English)
While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provenance for studying isolated sound-effect deepfakes. To support this direction, we present SynSFX, a large-scale corpus of 43374 clips (26452 synthetic, 16922 real) spanning 7 popular text-to-audio models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。