arXiv:2602.04702cs.SD2026-02中稿 · ICASSP 2026被引 3

通过精调帧级建模,提升语音伪造检测对细微篡改的识别能力。

Fine-Grained Frame Modeling in Multi-head Self-Attention for Speech Deepfake Detection

  • 用多头投票选关键帧,再跨层精修增强特征提取。
  • 在LA21、DF21、ITW数据集上EER分别降至0.90%、1.88%、6.64%。
  • 适合关注语音伪造检测中细节线索挖掘的研究者。

基于Transformer的模型在语音伪造检测中表现优异,主要得益于多头自注意力(MHSA)机制对帧级注意力得分的有效捕捉。由于语音伪造痕迹通常集中在时间维度的小范围区域,精细的帧级建模对于识别细微欺骗信号至关重要。本文提出细粒度帧建模(FGFM),首先通过多头投票(MHV)模块筛选最具信息量的帧,再经跨层精修(CLR)模块进一步优化,以增强模型对细微伪造线索的学习能力。实验表明,该方法优于基线模型,在LA21、DF21和ITW数据集上的等错误率(EER)分别为0.90%、1.88%和6.64%。多项基准测试的一致改进验证了细粒度建模在鲁棒语音伪造检测中的有效性。

原文摘要 · Abstract (English)

Transformer-based models have shown strong performance in speech deepfake detection, largely due to the effectiveness of the multi-head self-attention (MHSA) mechanism. MHSA provides frame-level attention scores, which are particularly valuable because deepfake artifacts often occur in small, localized regions along the temporal dimension of speech. This makes fine-grained frame modeling essential for accurately detecting subtle spoofing cues. In this work, we propose fine-grained frame modeling (FGFM) for MHSA-based speech deepfake detection, where the most informative frames are first selected through a multi-head voting (MHV) module. These selected frames are then refined via a cross-layer refinement (CLR) module to enhance the model's ability to learn subtle spoofing cues. Experimental results demonstrate that our method outperforms the baseline model and achieves Equal Error Rate (EER) of 0.90%, 1.88%, and 6.64% on the LA21, DF21, and ITW datasets, respectively. These consistent improvements across multiple benchmarks highlight the effectiveness of our fine-grained modeling for robust speech deepfake detection.

语音伪造自注意力细粒度建模深度伪造检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。