arXiv:2508.13320eess.AS2025-08

用少量样本快速适应新语音伪造,提升跨场景检测能力。

Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts

  • 基于自注意力原型网络实现少样本快速适应。
  • 仅需10个样本即实现32%相对错误率降低。
  • 适合应对未知合成方法与语言的语音伪造检测。

本文解决合成语音在分布偏移下的检测难题——即测试时出现训练中未见的合成方法、说话人、语言或音频条件。少样本学习通过少量分布内样本快速适应,是应对分布偏移的有力手段。我们提出一种自注意力原型网络,增强少样本适应的鲁棒性。为评估方法,系统比较传统零样本检测器与所提少样本检测器性能,严格控制训练条件以引入评估时的分布偏移。当分布偏移导致零样本性能下降时,所提少样本适应技术仅需10个分布内样本即可快速适应,在日语Deepfake数据集上实现高达32%的相对EER降低,在ASVspoof 2021 Deepfake数据集上实现20%的相对降低。

原文摘要 · Abstract (English)

We address the challenge of detecting synthesized speech under distribution shifts -- arising from unseen synthesis methods, speakers, languages, or audio conditions -- relative to the training data. Few-shot learning methods are a promising way to tackle distribution shifts by rapidly adapting on the basis of a few in-distribution samples. We propose a self-attentive prototypical network to enable more robust few-shot adaptation. To evaluate our approach, we systematically compare the performance of traditional zero-shot detectors and the proposed few-shot detectors, carefully controlling training conditions to introduce distribution shifts at evaluation time. In conditions where distribution shifts hamper the zero-shot performance, our proposed few-shot adaptation technique can quickly adapt using as few as 10 in-distribution samples -- achieving upto 32% relative EER reduction on deepfakes in Japanese language and 20% relative reduction on ASVspoof 2021 Deepfake dataset.

语音伪造少样本学习分布偏移检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。