arXiv:2502.14726cs.SDcs.CR2025-02被引 13

通过语音韵律特征检测音频深度伪造,准确率达93%且更抗攻击。

Pitch Imperfect: Detecting Audio Deepfakes Through Acoustic Prosodic Analysis

  • 基于六种经典韵律特征构建检测模型,聚焦人类识别语音的高层特征。
  • 在93%准确率下,误报率仅24.7%,且对对抗攻击有更强鲁棒性。
  • 可解释性强,能揭示影响判断的关键特征,适合安全与可信系统应用。

音频深度伪造正变得与真实语音难以区分,常骗过认证系统和人类听觉。现有方法多依赖低层音频特征或黑盒模型训练,而关注人类用于识别语音的高层特征可能更具长期稳健性。本文探索使用韵律——即人类语音的高层语言特征(如音高、语调、抖动)——作为检测音频深度伪造的基础方法。我们基于六种经典韵律特征构建检测器,实验表明其性能与社区主流基线模型相当,准确率达93%,等错误率(EER)为24.7%。更重要的是,通过采用$ L_{\infty} $范数对抗攻击测试,证明该模型比其他模型更具鲁棒性(其余模型准确率下降99.3%)。同时,利用注意力机制实现可解释性分析,识别出对决策影响最大的特征:抖动(Jitter)、闪烁(Shimmer)和平均基频(Mean Fundamental Frequency)。结果表明,以语言特征为基础的方法在性能相近的前提下,兼具更高鲁棒性和可解释性。

原文摘要 · Abstract (English)

Audio deepfakes are increasingly in-differentiable from organic speech, often fooling both authentication systems and human listeners. While many techniques use low-level audio features or optimization black-box model training, focusing on the features that humans use to recognize speech will likely be a more long-term robust approach to detection. We explore the use of prosody, or the high-level linguistic features of human speech (e.g., pitch, intonation, jitter) as a more foundational means of detecting audio deepfakes. We develop a detector based on six classical prosodic features and demonstrate that our model performs as well as other baseline models used by the community to detect audio deepfakes with an accuracy of 93% and an EER of 24.7%. More importantly, we demonstrate the benefits of using a linguistic features-based approach over existing models by applying an adaptive adversary using an $L_{\infty}$ norm attack against the detectors and using attention mechanisms in our training for explainability. We show that we can explain the prosodic features that have highest impact on the model's decision (Jitter, Shimmer and Mean Fundamental Frequency) and that other models are extremely susceptible to simple $L_{\infty}$ norm attacks (99.3% relative degradation in accuracy). While overall performance may be similar, we illustrate the robustness and explainability benefits to a prosody feature approach to audio deepfake detection.

音频伪造韵律分析可解释性对抗鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。