首个面向东南亚语言的语音伪造检测基准,提出轻量级模型提升实际应用能力。
Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

- 构建跨语种、多说话人、多编码器的语音伪造检测基准SEA-CF
- 现有大模型在东南亚语音上表现差,因语言特异性导致泛化失败
- 提出轻量级模型GARUDA,兼顾性能与实用性,适合低资源场景
语音伪造(Codecfakes, CFs)通过音频语言模型(ALMs)生成,以神经音频编码器(NACs)为核心机制。此类伪造语音分布特征不同于基于声码器的伪造,导致现有检测器在训练数据为声码器时难以泛化至CFs。尽管已有相关检测基准,但主要局限于英语和部分中文,东南亚(SEA)语言尚无系统研究。为此,本文首次构建了覆盖多种东南亚语言、多样说话人及广泛NAC架构的大规模基准SEA-CF,基于公开真实语音语料合成。实验表明,基于英语数据训练的先进检测器在东南亚语音上表现显著下降,归因于语言特有的音系结构、声调差异与丰富韵律特征。我们进一步对近期SOTA ALMs进行零样本与微调评估,发现微调可提升性能,但模型过大,不适用于低资源与低延迟场景。为此,我们提出新型小型音频语言模型GARUDA,专为伪造检测设计,在保持轻量化的同时实现优异性能。大量实验证明,该小模型优于主流端到端及基于ALM的基线,为东南亚语言乃至更广范围的鲁棒语音伪造检测提供了新路径。
原文摘要 · Abstract (English)
Codecfakes (CFs) are a type of speech deepfakes generated through Audio Language Models (ALMs), with Neural Audio Codecs (NACs) forming the core mechanism for speech encoding and generation. CFs exhibit distributional characteristics that differ from vocoder-based deepfakes, causing detectors trained on vocoder data to generalize poorly to CFs detection. Although this has led to the development of CF detection benchmarks, existing resources are largely confined to English -- and to a limited extent Chinese -- leaving South-East Asian (SEA) languages unexplored. To bridge this gap, we introduce SEA-CF, the first large-scale benchmark for CF detection spanning multiple SEA languages, diverse speaker profiles, and a wide range of NAC architectures. SEA-CF is constructed by synthesizing publicly available real speech corpora. Our experiments show that state-of-the-art (SOTA) CF detectors trained on English-centric datasets fail to generalize to SEA speech due to language-specific phonetic structures, tonal variations, and rich prosodic diversity. We further conduct a comprehensive zero-shot and fine-tuned evaluation of recent SOTA ALMs on SEA-CF. Fine-tuning the ALMs improves performance, however, these are very large being impractical for real-world application due to their scale, particularly in low-resource and latency-constrained settings. To address this limitation, we propose a novel small-ALM, GARUDA tailored for CF detection, which delivers strong performance while remaining lightweight. Extensive evaluations demonstrate that the proposed Small-ALM outperforms strong end-to-end and ALM-based baselines, establishing a new, practical direction for robust CF detection in SEA languages and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。