首个面向东南亚语种的语音伪造检测数据集,解决跨语言检测失效问题。
SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
- 构建涵盖6种东南亚语言的300+小时真实与伪造语音配对数据集。
- 跨语言检测性能下降明显,但在该数据集上微调后显著恢复。
- 适合语音安全、多语种检测方向的研究者使用。
东南亚数字经济发展迅速,语音伪造风险加剧,但现有数据集对东南亚语言覆盖稀疏,导致模型在该区域表现不佳。由于合成质量差异、语言特性不同及数据稀缺,基于高资源语言训练的检测模型在东南亚语种上严重失效。为此,我们提出SEA-Spoof,首个专为东南亚语言设计的大规模语音伪造检测数据集。该数据集包含300+小时的泰米尔语、印地语、泰语、印尼语、马来语和越南语的真实与伪造语音配对样本,伪造样本来自多种先进开源与商用系统,涵盖广泛风格与保真度。基准测试显示,现有检测模型存在严重跨语言退化,但在SEA-Spoof上微调后,各语言及合成源下的性能均显著提升。结果凸显了针对东南亚语种研究的紧迫性,并确立SEA-Spoof作为构建鲁棒、跨语言、抗欺诈检测系统的基石。
原文摘要 · Abstract (English)
The rapid growth of the digital economy in South-East Asia (SEA) has amplified the risks of audio deepfakes, yet current datasets cover SEA languages only sparsely, leaving models poorly equipped to handle this critical region. This omission is critical: detection models trained on high-resource languages collapse when applied to SEA, due to mismatches in synthesis quality, language-specific characteristics, and data scarcity. To close this gap, we present SEA-Spoof, the first large-scale Audio Deepfake Detection (ADD) dataset especially for SEA languages. SEA-Spoof spans 300+ hours of paired real and spoof speech across Tamil, Hindi, Thai, Indonesian, Malay, and Vietnamese. Spoof samples are generated from a diverse mix of state-of-the-art open-source and commercial systems, capturing wide variability in style and fidelity. Benchmarking state-of-the-art detection models reveals severe cross-lingual degradation, but fine-tuning on SEA-Spoof dramatically restores performance across languages and synthesis sources. These results highlight the urgent need for SEA-focused research and establish SEA-Spoof as a foundation for developing robust, cross-lingual, and fraud-resilient detection systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。