arXiv:2603.16914cs.SDcs.AI2026-03被引 1

利用量化层级结构提升语音伪造检测精度

Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection

  • 通过可学习权重建模多级量化贡献,捕捉伪造痕迹
  • 在ASVspoof2019上相对误报率降低46.2%
  • 仅更新4.4%参数,适合轻量级部署

神经音频编解码器通过残差向量量化(RVQ)对语音进行离散化,形成从粗到细的量化层级。尽管编解码模型已用于表征学习,但其离散结构在语音伪造检测中仍未被充分挖掘。不同量化层级捕获互补声学线索:早期量化器编码粗略结构,后期量化器细化残留细节,揭示合成伪影。现有系统或依赖连续编码特征,或忽略量化层级。我们提出一种层次感知表示学习框架,通过可学习全局权重建模量化层级贡献,生成与取证线索对齐的结构化编解码表示。保持语音编码器主干冻结,仅更新4.4%额外参数,方法在ASVspoof 2019上实现46.2%的相对等错误率(EER)下降,在ASVspoof5上实现13.9%的相对下降,优于强基线。

原文摘要 · Abstract (English)

Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains underutilized in speech deepfake detection. In particular, different quantization levels capture complementary acoustic cues, where early quantizers encode coarse structure and later quantizers refine residual details that reveal synthesis artifacts. Existing systems either rely on continuous encoder features or ignore this quantizer-level hierarchy. We propose a hierarchy-aware representation learning framework that models quantizer-level contributions through learnable global weighting, enabling structured codec representations aligned with forensic cues. Keeping the speech encoder backbone frozen and updating only 4.4% additional parameters, our method achieves relative EER reductions of 46.2% on ASVspoof 2019 and 13.9% on ASVspoof5 over strong baselines.

语音伪造检测编解码器量化层级取证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。