arXiv:2509.21728cs.SD2025-09被引 1

无需训练即可检测新型语音伪造,靠检索和声纹匹配实现。

Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching

  • 通过检索已有知识库+声纹匹配,不需重新训练
  • 在DeepFake-Eval-2024上性能媲美有监督模型
  • 适合快速响应新出现的语音伪造攻击

基于大模型和大规模数据集的现代语音伪造检测器虽表现良好,但在面对零日攻击(即由全新合成方法生成的音频)时效果下降。传统微调方法响应慢,难以满足即时需求。本文提出一种无训练的检索增强型检测框架,利用知识表示与声纹匹配技术。该框架采用简单有效的检索与集成策略,在DeepFake-Eval-2024基准上达到与监督基线及其微调版本相当的性能,且无需额外训练。我们还对声纹属性进行了消融实验,并通过简单的无训练融合策略验证了框架在跨数据库场景下的泛化能力。

原文摘要 · Abstract (English)

Modern audio deepfake detectors built on foundation models and large training datasets achieve promising detection performance. However, they struggle with zero-day attacks, where the audio samples are generated by novel synthesis methods that models have not seen from reigning training data. Conventional approaches fine-tune the detector, which can be problematic when prompt response is needed. This paper proposes a training-free retrieval-augmented framework for zero-day audio deepfake detection that leverages knowledge representations and voice profile matching. Within this framework, we propose simple yet effective retrieval and ensemble methods that reach performance comparable to supervised baselines and their fine-tuned counterparts on the DeepFake-Eval-2024 benchmark, without any additional model training. We also conduct ablation on voice profile attributes, and demonstrate the cross-database generalizability of the framework with introducing simple and training-free fusion strategies.

语音伪造零日检测检索增强声纹匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。