arXiv:2505.13976eess.AScs.SD2025-05中稿 · Interspeech 2025被引 2

用语音自然度指导训练,提升语音伪造检测鲁棒性

Naturalness-Aware Curriculum Learning with Dynamic Temperature for Speech Deepfake Detection

  • 结合真实标签与主观评分动态评估样本难度,渐进式训练
  • 在ASVspoof 2021上实现EER降低23%,无需改模型结构
  • 适合关注语音伪造检测泛化能力的研究者和安全应用

近年来,语音伪造检测(SDD)在基于伪影的检测方面取得显著进展,但多数模型忽略了语音自然度这一关键判别线索。本文提出一种自然度感知的课程学习框架,利用语音自然度提升SDD模型的鲁棒性与泛化能力。该方法结合真实标签与平均意见分(MOS)衡量样本难度,动态调整训练顺序,逐步引入更难样本。为进一步增强泛化性,引入基于语音自然度的动态温度缩放机制。在ASVspoof 2021 DF数据集上的实验表明,该方法在不修改模型架构的前提下,实现了相对EER降低23%。消融实验证实了自然度感知训练策略对SDD任务的有效性。

原文摘要 · Abstract (English)

Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed speech. This study proposes naturalness-aware curriculum learning, a novel training framework that leverages speech naturalness to enhance the robustness and generalization of SDD. This approach measures sample difficulty using both ground-truth labels and mean opinion scores, and adjusts the training schedule to progressively introduce more challenging samples. To further improve generalization, a dynamic temperature scaling method based on speech naturalness is incorporated into the training process. A 23% relative reduction in the EER was achieved in the experiments on the ASVspoof 2021 DF dataset, without modifying the model architecture. Ablation studies confirmed the effectiveness of naturalness-aware training strategies for SDD tasks.

语音伪造检测课程学习自然度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。