arXiv:2509.20679cs.SDcs.AI2025-09中稿 · INTERSPEECH 2026被引 1

用多个质量感知中心提升语音伪造检测精度

QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection

  • 引入多质量中心建模真实语音的多样性
  • 在野数据集上等错误率降至5.09%
  • 无需标注质量即可推理,适合实际部署

近期研究发现,通过单个中心点建模真实语音的紧凑分布,可检测未见过的语音伪造攻击。然而,单中心假设会简化真实语音表示,忽略如语音质量等重要线索,而语音质量可通过现有模型估算为均值意见分数(MOS)。本文提出QAMO:一种质量感知的多中心一分类学习方法,用于语音伪造检测。QAMO通过引入多个质量感知中心,使每个中心代表特定语音质量子空间,从而更好捕捉真实语音内部变异性。此外,该方法支持多中心集成评分策略,优化决策阈值并减少推理时对质量标签的依赖。使用两个中心分别表示高质量与低质量语音,所提QAMO在In-the-Wild数据集上实现5.09%的等错误率,优于先前的一类学习和质量感知系统。

原文摘要 · Abstract (English)

Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-centroid assumption can oversimplify the bona fide speech representation and overlook useful cues, such as speech quality, which reflects the naturalness of the speech. Speech quality can be easily obtained using existing speech quality assessment models that estimate it through Mean Opinion Score. In this paper, we propose QAMO: Quality-Aware Multi-Centroid One-Class Learning for speech deepfake detection. QAMO extends conventional one-class learning by introducing multiple quality-aware centroids. In QAMO, each centroid is optimized to represent a distinct speech quality subspaces, enabling better modeling of intra-class variability in bona fide speech. In addition, QAMO supports a multi-centroid ensemble scoring strategy, which improves decision thresholding and reduces the need for quality labels during inference. With two centroids to represent high- and low-quality speech, our proposed QAMO achieves an equal error rate of 5.09% in In-the-Wild dataset, outperforming previous one-class and quality-aware systems.

语音伪造检测一分类学习质量感知多中心建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。