arXiv:2509.21676eess.AScs.LG2025-09被引 2

通过语音韵律感知提升对情感合成语音的欺骗检测能力

HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech

  • 分两阶段训练,利用韵律特征增强模型对自然语音的感知
  • 在跨语言、情感等挑战性数据集上超越现有基线
  • 适合需要抵御高级语音伪造攻击的系统开发者

当前的反欺骗系统对富有表现力和情感的合成语音仍易受攻击,因其很少利用韵律作为判别线索。韵律是人类表达与情感的核心,人类会本能依赖基频(F0)模式和有声/无声结构来区分自然与合成语音。本文提出HuLA,一种两阶段的韵律感知多任务学习框架用于欺骗检测。第一阶段,使用自监督学习(SSL)骨干网络在真实语音上训练,辅以基频预测和有声/无声分类任务,增强其对自然韵律变化的捕捉能力,类比人类感知学习。第二阶段,模型在真实与合成数据上联合优化欺骗检测与韵律任务,利用韵律感知识别自然语音与富有表现力的合成语音之间的不匹配。实验表明,HuLA在具有挑战性的域外数据集上持续优于强基线,包括表现力、情感及跨语言攻击场景。结果证明,显式韵律监督结合SSL嵌入能显著提升对先进合成语音攻击的鲁棒性。

原文摘要 · Abstract (English)

Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans instinctively use prosodic cues such as F0 patterns and voiced/unvoiced structure to distinguish natural from synthetic speech. In this paper, we propose HuLA, a two-stage prosody-aware multi-task learning framework for spoof detection. In Stage 1, a self-supervised learning (SSL) backbone is trained on real speech with auxiliary tasks of F0 prediction and voiced/unvoiced classification, enhancing its ability to capture natural prosodic variation similar to human perceptual learning. In Stage 2, the model is jointly optimized for spoof detection and prosody tasks on both real and synthetic data, leveraging prosodic awareness to detect mismatches between natural and expressive synthetic speech. Experiments show that HuLA consistently outperforms strong baselines on challenging out-of-domain dataset, including expressive, emotional, and cross-lingual attacks. These results demonstrate that explicit prosodic supervision, combined with SSL embeddings, substantially improves robustness against advanced synthetic speech attacks.

语音安全韵律感知反欺骗多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。