首个大规模西班牙语语音伪造检测数据集,助力识别合成语音。
HISPASpoof: A New Dataset For Spanish Speech Forensics
- 构建包含六种口音的西班牙语真实与合成语音数据集
- 英文检测模型在西班牙语上效果差,用该数据集训练可显著提升性能
- 适用于语音安全、跨语言检测研究者
零样本语音克隆(VC)和文本转语音(TTS)技术快速发展,生成高度逼真的合成语音,引发滥用担忧。尽管已有大量针对英语和中文的检测方法,但全球超6亿人使用的西班牙语在语音取证领域仍被严重忽视。为此,我们推出HISPASpoof——首个大规模西班牙语合成语音检测与溯源数据集。该数据集包含来自六个口音的真实语音(源自公开语料库)及六种零样本TTS系统生成的合成语音。我们评估了五种代表性检测方法,发现基于英语训练的模型无法泛化至西班牙语,而使用HISPASpoof训练则能显著提升检测准确率。同时,我们还评估了合成语音溯源性能,即识别合成语音的生成方法。HISPASpoof为推进可靠且包容的西班牙语语音取证提供了关键基准。
原文摘要 · Abstract (English)
Zero-shot Voice Cloning (VC) and Text-to-Speech (TTS) methods have advanced rapidly, enabling the generation of highly realistic synthetic speech and raising serious concerns about their misuse. While numerous detectors have been developed for English and Chinese, Spanish-spoken by over 600 million people worldwide-remains underrepresented in speech forensics. To address this gap, we introduce HISPASpoof, the first large-scale Spanish dataset designed for synthetic speech detection and attribution. It includes real speech from public corpora across six accents and synthetic speech generated with six zero-shot TTS systems. We evaluate five representative methods, showing that detectors trained on English fail to generalize to Spanish, while training on HISPASpoof substantially improves detection. We also evaluate synthetic speech attribution performance on HISPASpoof, i.e., identifying the generation method of synthetic speech. HISPASpoof thus provides a critical benchmark for advancing reliable and inclusive speech forensics in Spanish.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。