arXiv:2601.19573eess.AS2026-01中稿 · ICASSP 2026被引 1

检测对话开头的短语音伪造,提升真实场景下防骗能力

Audio Deepfake Detection at the First Greeting: "Hi!"

  • 设计轻量模块增强短时低质语音的时频特征辨识力
  • 在0.5-2秒极短语音上实现超90%检测准确率且抗干扰强
  • 适合部署于手机等边缘设备,实时防护语音诈骗

本文聚焦真实通信环境下极短语音(0.5-2.0秒)的音频深度伪造检测,目标是识别对话起始处的合成语音,如骗子说的“Hi”。提出S-MGAA,一种轻量级多粒度自适应时频注意力扩展模型,用于提升短时、受损输入的判别性表征学习能力。S-MGAA集成两个定制模块:像素-通道增强模块(PCEM)放大细粒度时频显著性,频率补偿增强模块(FCEM)通过多尺度频率建模与自适应频时交互补充有限时间证据。大量实验表明,S-MGAA持续优于九个先进基线,在降噪、压缩等退化条件下保持强鲁棒性,并具备低实时因子(RTF)、合理浮点运算量(GFLOPs)、紧凑参数量和低训练成本,展现其在通信系统与边缘设备中实时部署的强大潜力。

原文摘要 · Abstract (English)

This paper focuses on audio deepfake detection under real-world communication degradations, with an emphasis on ultra-short inputs (0.5-2.0s), targeting the capability to detect synthetic speech at a conversation opening, e.g., when a scammer says "Hi." We propose Short-MGAA (S-MGAA), a novel lightweight extension of Multi-Granularity Adaptive Time-Frequency Attention, designed to enhance discriminative representation learning for short, degraded inputs subjected to communication processing and perturbations. The S-MGAA integrates two tailored modules: a Pixel-Channel Enhanced Module (PCEM) that amplifies fine-grained time-frequency saliency, and a Frequency Compensation Enhanced Module (FCEM) to supplement limited temporal evidence via multi-scale frequency modeling and adaptive frequency-temporal interaction. Extensive experiments demonstrate that S-MGAA consistently surpasses nine state-of-the-art baselines while achieving strong robustness to degradations and favorable efficiency-accuracy trade-offs, including low RTF, competitive GFLOPs, compact parameters, and reduced training cost, highlighting its strong potential for real-time deployment in communication systems and edge devices.

语音伪造检测轻量化模型边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。