arXiv:2505.23962eess.AScs.LG2025-05中稿 · Interspeech 2025被引 8

情绪化合成语音正威胁防欺骗系统,新数据集与模型提升防御能力

Can Emotion Fool Anti-spoofing?

  • 构建情绪化语音合成数据集EmoSpoof-TTS,揭示情感对欺骗检测的影响
  • 现有模型在情绪化合成语音上准确率下降超30%,表现受情绪类型影响显著
  • 提出门控集成模型GEM,跨情绪场景统一提升防欺骗性能

传统反欺骗方法依赖以中性情绪为主的合成语音数据集和模型,忽视了真实场景中多样的情感变化。这导致其对高质量、富有情感的合成语音的鲁棒性未知。为此,我们提出了EmoSpoof-TTS数据集,包含多种情绪的文本转语音样本。分析显示,现有反欺骗模型在情感化合成语音上表现明显下降,存在被针对性攻击的风险。即使在情感数据上训练,模型仍因未充分关注情感特征而表现不均,不同情绪间性能差异超过30%。这表明需建立以情感为核心的反欺骗范式。我们提出GEM模型,采用情感识别门控的专家集成结构,能有效应对所有情绪及中性状态,显著提升防御能力。EmoSpoof-TTS数据集已公开:https://emospoof-tts.github.io/Dataset/

原文摘要 · Abstract (English)

Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness against high-quality, emotionally expressive synthetic speech is uncertain. We address this by introducing EmoSpoof-TTS, a corpus of emotional text-to-speech samples. Our analysis shows existing anti-spoofing models struggle with emotional synthetic speech, exposing risks of emotion-targeted attacks. Even trained on emotional data, the models underperform due to limited focus on emotional aspect and show performance disparities across emotions. This highlights the need for emotion-focused anti-spoofing paradigm in both dataset and methodology. We propose GEM, a gated ensemble of emotion-specialized models with a speech emotion recognition gating network. GEM performs effectively across all emotions and neutral state, improving defenses against spoofing attacks. We release the EmoSpoof-TTS Dataset: https://emospoof-tts.github.io/Dataset/

语音合成反欺骗情感识别安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。