arXiv:2506.11532eess.AScs.SD2025-06中稿 · Interspeech 2025被引 6

用尖锐度理论提升语音伪造检测的泛化能力

From Sharpness to Better Generalization for Speech Deepfake Detection

  • 以模型尖锐度为理论指标,分析其在跨域场景下的变化规律
  • 采用SAM方法降低尖锐度,使模型在未见数据上表现更稳定
  • 证实尖锐度与泛化性能显著相关,适合提升鲁棒性研究者参考

语音伪造检测(SDD)的泛化能力仍是关键挑战。现有方法多关注鲁棒性提升,但缺乏解释模型性能的理论框架。本文将尖锐度作为泛化性的理论代理,分析其在域偏移下的响应:发现尖锐度在未见条件下上升,表明模型敏感性增强。基于此,应用尖锐度感知最小化(SAM)显式降低尖锐度,实现跨多样未见测试集更优且稳定的性能。进一步相关性分析表明,在多数测试设置中,尖锐度与泛化能力存在统计显著关系。结果表明,尖锐度可作为SDD中泛化性的理论指标,尖锐度感知训练是提升鲁棒性的有效策略。

原文摘要 · Abstract (English)

Generalization remains a critical challenge in speech deepfake detection (SDD). While various approaches aim to improve robustness, generalization is typically assessed through performance metrics like equal error rate without a theoretical framework to explain model performance. This work investigates sharpness as a theoretical proxy for generalization in SDD. We analyze how sharpness responds to domain shifts and find it increases in unseen conditions, indicating higher model sensitivity. Based on this, we apply Sharpness-Aware Minimization (SAM) to reduce sharpness explicitly, leading to better and more stable performance across diverse unseen test sets. Furthermore, correlation analysis confirms a statistically significant relationship between sharpness and generalization in most test settings. These findings suggest that sharpness can serve as a theoretical indicator for generalization in SDD and that sharpness-aware training offers a promising strategy for improving robustness.

语音伪造泛化能力尖锐度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。