arXiv:2605.03079cs.SDcs.LG2026-05被引 1

从音素层面分析情绪化语音伪造,提升检测可解释性。

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

论文配图:Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
图 1 · 摘自论文原文
  • 基于音素对齐的WavLM嵌入,逐音素分析情绪化语音
  • 复杂音素(如元音、擦音)差异更大,更易被检测
  • 适用于多情绪、多合成系统,适合语音安全研究者

情感语音转换(EVC)技术的进步使得生成富有表现力的合成语音成为可能,也带来了新的音频深度伪造检测挑战。现有方法将语音视为均质信号,忽视其内部音素结构,限制了在情绪化条件下的可解释性。本文提出一种音素级框架,利用真实与EVC生成语音在相同情绪条件下、共享转录文本、音素对齐的TextGrid和基于WavLM的嵌入进行分析。结果表明,音素行为随类别而异,复杂元音和擦音表现出更高差异,而简单音素则更稳定。分布差异较大的音素在多种情绪和合成系统下均更易被检测。这些发现证明,音素级分析是检测情绪化伪造语音的有效且可解释的方法。

原文摘要 · Abstract (English)

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook its internal phonetic structure, limiting their interpretability in emotionally conditioned settings. In this work, we propose a phoneme-level framework to analyze emotionally manipulated synthetic speech using real and EVC-generated speech under matched emotional conditions with shared transcripts, phoneme-aligned TextGrids, and WavLM-based embeddings. Our results show that phoneme behavior varies across categories, with complex vowels and fricatives exhibiting higher divergence while simpler phonemes remain more stable. Phonemes with larger distributional differences are also found to be more easily detected, consistently across multiple emotions and synthesis systems. These findings demonstrate that phoneme-level analysis is an effective and interpretable approach for detecting emotionally manipulated synthetic speech.

语音伪造音素分析自监督学习情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。