arXiv:2507.08626cs.SDeess.AS2025-07ICCV被引 10

通过音素级分析提升说话人伪造检测的可解释性与鲁棒性

Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection

  • 将参考音频分解为音素,构建精细说话人特征档案
  • 测试时逐音素比对,实现细粒度伪造痕迹识别
  • 适合需要可解释性的多媒体鉴伪场景

生成式AI的进展使语音伪造广泛可及,严重威胁数字信任。为应对这一挑战,现有研究提出了关注特定个体的说话人-兴趣点(POI)检测方法,通过建模和分析其独特声学特征来识别模仿行为。尽管性能优异,但现有方法缺乏粒度和可解释性。本文提出一种基于POI的语音伪造检测新方法,实现音素级别分析:将参考音频分解为音素,构建详细说话人特征档案;推理时,对测试样本的每个音素逐个与档案比对,实现细粒度合成伪迹检测。该方法在准确率上与传统方法相当,同时具备更优的鲁棒性和可解释性,是多媒体鉴伪中可解释、以说话人为中心的新型检测方向。

原文摘要 · Abstract (English)

Recent advances in generative AI have made the creation of speech deepfakes widely accessible, posing serious challenges to digital trust. To counter this, various speech deepfake detection strategies have been proposed, including Person-of-Interest (POI) approaches, which focus on identifying impersonations of specific individuals by modeling and analyzing their unique vocal traits. Despite their excellent performance, the existing methods offer limited granularity and lack interpretability. In this work, we propose a POI-based speech deepfake detection method that operates at the phoneme level. Our approach decomposes reference audio into phonemes to construct a detailed speaker profile. In inference, phonemes from a test sample are individually compared against this profile, enabling fine-grained detection of synthetic artifacts. The proposed method achieves comparable accuracy to traditional approaches while offering superior robustness and interpretability, key aspects in multimedia forensics. By focusing on phoneme analysis, this work explores a novel direction for explainable, speaker-centric deepfake detection.

语音伪造可解释性音素分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。