arXiv:2505.23619cs.SDcs.LG2025-05被引 3

用高斯过程实现少量语音数据下的深度伪造检测自适应。

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

  • 基于高斯过程构建可快速适应新语音模型的检测框架。
  • 仅需少量样本即可在新生成模型上保持高准确率。
  • 适合需要快速部署、应对新型语音伪造的技术团队。

文本到语音(TTS)模型,尤其是语音克隆技术的快速发展,推动了对可适应、高效深度伪造检测方法的需求。随着TTS系统持续演进,检测模型必须能在仅有极少数据的情况下,有效适应未见过的生成模型。本文提出ADD-GP,一种基于高斯过程(GP)分类器的音频深度伪造检测(ADD)少样本自适应框架。我们展示了强大深度嵌入模型与高斯过程灵活性相结合,可在性能和适应性上取得优异表现。此外,该方法还可用于个性化检测,对新型TTS模型更具鲁棒性,并支持单样本自适应。为支持评估,我们构建了一个基准数据集,采用最新语音克隆模型生成数据。

原文摘要 · Abstract (English)

Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS systems continue to evolve, detection models must be able to efficiently adapt to previously unseen generation models with minimal data. This paper introduces ADD-GP, a few-shot adaptive framework based on a Gaussian Process (GP) classifier for Audio Deepfake Detection (ADD). We show how the combination of a powerful deep embedding model with the Gaussian processes flexibility can achieve strong performance and adaptability. Additionally, we show this approach can also be used for personalized detection, with greater robustness to new TTS models and one-shot adaptability. To support our evaluation, a benchmark dataset is constructed for this task using new state-of-the-art voice cloning models.

语音伪造少样本学习高斯过程检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。