arXiv:2604.13229eess.AS2026-04被引 2

通过语音韵律特征学习,提升语音伪造检测对情感攻击的鲁棒性。

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks

  • 基于音高、语音活动和能量的韵律掩码预测,增强模型对真实语音变异的建模。
  • 在ASVspoof 2024上将误报率从39.62%降至7.38%,情感攻击场景下相对降低50%。
  • 适合关注语音安全与对抗情感欺骗的系统设计者与研究者。

语音深度伪造检测(SDD)系统在标准基准数据集上表现良好,但在表达丰富和情绪化伪造攻击下泛化能力差。许多方法依赖带有伪造样本的数据训练,学习的是数据集特有的伪影而非可迁移的真实语音特征。相比之下,人类能内化真实语音的多样性,并将其作为识别伪造的依据。本文提出ProSDD,一种两阶段框架:第一阶段通过监督式掩码预测,学习说话人相关的韵律变化(基于音高、语音活动和能量);第二阶段联合优化该目标与伪造分类任务。ProSDD在ASVspoof 2019和2024训练下均显著优于基线,在2024训练条件下将EER从39.62%降至7.38%,在情感伪造数据集EmoFake和EmoSpoof-TTS上实现50%的相对性能提升。

原文摘要 · Abstract (English)

Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific artifacts rather than transferable cues of natural speech. In contrast, humans internalize variability in real speech and detect fakes as deviations from it. We introduce ProSDD, a two-stage framework that enriches model embeddings through supervised masked prediction of speaker-conditioned prosodic variation based on pitch, voice activity, and energy. Stage I learns prosodic variability from real speech, and Stage II jointly optimizes this objective with spoof classification. ProSDD consistently outperforms baselines under both ASVspoof 2019 and 2024 training, reducing ASVspoof 2024 EER from 25.43% to 16.14% (2019-trained) and from 39.62% to 7.38% (2024-trained), while achieving 50% relative reductions on EmoFake and EmoSpoof-TTS.

语音伪造韵律建模对抗攻击深度伪造检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。