arXiv:2412.12619cs.SDcs.AI2024-12被引 11

通过检测音素级特征不一致,有效识别高保真语音深度伪造

Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes

  • 设计自适应音素池化,提取每段语音特有的音素级特征
  • 在4个基准数据集上超越现有方法,准确率显著提升
  • 适合语音安全、内容审核等需要防伪的应用场景

近年来,文本转语音和语音转换技术的发展使得合成语音愈发逼真。尽管这些创新带来诸多便利,但若被恶意利用,将引发严重安全问题。因此,迫切需要检测合成语音信号。音素特征为深度伪造检测提供了强大表征。然而,以往基于音素的方法多聚焦特定音素,忽视了整个音素序列中的时序不一致性。本文提出一种新机制,通过识别音素级语音特征的不一致性来检测语音深度伪造。我们设计了一种自适应音素池化技术,从帧级语音数据中提取样本特定的音素级特征。在预训练音频模型提取的特征基础上,应用于未见过的深度伪造数据集,发现伪造样本普遍存在音素级不一致性。为进一步提升检测精度,我们提出使用图注意力网络建模音素级特征的时序依赖关系,并引入随机音素替换增强技术以增加训练时的特征多样性。在四个基准数据集上的大量实验表明,该方法优于现有最先进检测方法。

原文摘要 · Abstract (English)

Recent advancements in text-to-speech and speech conversion technologies have enabled the creation of highly convincing synthetic speech. While these innovations offer numerous practical benefits, they also cause significant security challenges when maliciously misused. Therefore, there is an urgent need to detect these synthetic speech signals. Phoneme features provide a powerful speech representation for deepfake detection. However, previous phoneme-based detection approaches typically focused on specific phonemes, overlooking temporal inconsistencies across the entire phoneme sequence. In this paper, we develop a new mechanism for detecting speech deepfakes by identifying the inconsistencies of phoneme-level speech features. We design an adaptive phoneme pooling technique that extracts sample-specific phoneme-level features from frame-level speech data. By applying this technique to features extracted by pre-trained audio models on previously unseen deepfake datasets, we demonstrate that deepfake samples often exhibit phoneme-level inconsistencies when compared to genuine speech. To further enhance detection accuracy, we propose a deepfake detector that uses a graph attention network to model the temporal dependencies of phoneme-level features. Additionally, we introduce a random phoneme substitution augmentation technique to increase feature diversity during training. Extensive experiments on four benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art detection methods.

语音伪造深度学习音素特征安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。