arXiv:2501.08238cs.SDeess.AS2025-01中稿 · TASLP 2026被引 9

构建首个大规模编解码器伪造语音数据集,助力识别新型深度伪造语音。

CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech

  • 用31种开源编解码器重合成语音作为检测训练数据
  • 验证频域解码器与解耦辅助目标提升检测效果
  • 提出编解码器分类新框架,适配细粒度检测研究

随着神经音频编解码器的快速发展,基于编解码器的语音生成(CoSG)系统日益强大。然而,这也催生了高度逼真的深度伪造语音,使个人声音模仿和虚假信息传播更加容易。我们将此类由CoSG系统生成的伪造语音称为CodecFake。检测CodecFake是一项紧迫挑战,但现有系统多聚焦于传统语音合成模型生成的假语音。本文提出CodecFake+,一个大规模数据集,旨在推动CodecFake检测研究。据我们所知,CodecFake+是涵盖最多样编解码器架构的最大的数据集。训练集通过31个公开开源编解码器进行重合成生成,评估集包含来自17个先进CoSG模型的网络采集数据。我们还提出了一个综合分类体系,按向量量化器、辅助目标和解码器类型对编解码器进行分类。该数据集与分类体系支持多层次分析,揭示成功检测的关键因素。在单个编解码器层面,我们验证了使用编解码器重合成语音(CoRS)作为训练数据的有效性;在分类层面,发现包含解耦辅助目标或频域解码器的重合成模型检测性能最佳。此外,综合所有CoRS数据来看,我们的分类体系可用于筛选更优训练数据以提升检测性能。总体而言,我们期望CodecFake+能为通用与细粒度反欺骗模型开发提供重要资源。

原文摘要 · Abstract (English)

With the rapid advancement of neural audio codecs, codec-based speech generation (CoSG) systems have become highly powerful. Unfortunately, CoSG also enables the creation of highly realistic deepfake speech, making it easier to mimic an individual's voice and spread misinformation. We refer to this emerging deepfake speech generated by CoSG systems as CodecFake. Detecting such CodecFake is an urgent challenge, yet most existing systems primarily focus on detecting fake speech generated by traditional speech synthesis models. In this paper, we introduce CodecFake+, a large-scale dataset designed to advance CodecFake detection. To our knowledge, CodecFake+ is the largest dataset encompassing the most diverse range of codec architectures. The training set is generated through re-synthesis using 31 publicly available open-source codec models, while the evaluation set includes web-sourced data from 17 advanced CoSG models. We also propose a comprehensive taxonomy that categorizes codecs by their root components: vector quantizer, auxiliary objectives, and decoder types. Our proposed dataset and taxonomy enable detailed analysis at multiple levels to discern the key factors for successful CodecFake detection. At the individual codec level, we validate the effectiveness of using codec re-synthesized speech (CoRS) as training data for large-scale CodecFake detection. At the taxonomy level, we show that detection performance is strongest when the re-synthesis model incorporates disentanglement auxiliary objectives or a frequency-domain decoder. Furthermore, from the perspective of using all the CoRS training data, we show that our proposed taxonomy can be used to select better training data for improving detection performance. Overall, we envision that CodecFake+ will be a valuable resource for both general and fine-grained exploration to develop better anti-spoofing models against CodecFake.

语音伪造数据集检测编解码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。