arXiv:2509.24674eess.AS2025-09中稿 · Odyssey 2026被引 3

提出零样本语音伪造溯源新框架,区分分布内与分布外场景表现更优。

Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing

  • 基于说话人验证思路,结合AAM损失与RegMixup增强特征表达。
  • 分布内场景下少样本模型误差率低至13.11%,分布外场景零样本更胜一筹。
  • 适用于未知伪造源的跨域溯源,适合安全检测与反欺诈研究者。

我们提出一种受说话人验证启发的新型零样本源溯源框架。将SSL-AASIST用于攻击分类,通过AAM损失和RegMixup增强嵌入表示,并确保训练攻击与指纹-试验对中的攻击互不重叠。在后端评分阶段,探索了零样本(余弦相似度、孪生网络)和少样本(MLP、孪生网络)方法。在最新提出的STOPA数据集上,开放集设置下的实验表明:在分布内(ID)场景中,少样本学习更具优势;而在分布外(OOD)场景中,零样本方法表现更佳。具体而言,在分布内试验中,少样本孪生网络与MLP的等错误率(EER)分别为17.72%和13.11%,显著优于零样本余弦相似度的29.91%;而在分布外试验中,零样本余弦相似度达到16.43%,优于少样本孪生网络的23.47%和MLP的21.57%。

原文摘要 · Abstract (English)

We propose a novel zero-shot source tracing framework inspired by speaker verification. We adapt SSL-AASIST for attack classification, enhancing embeddings with AAM loss and RegMixup, and ensure that training attacks are disjoint from those forming fingerprint-trial pairs. For backend scoring in attack verification, we explore both zero-shot approaches (cosine similarity and Siamese) and few-shot approaches (MLP and Siamese). Experiments on our recently introduced STOPA dataset with an open set setting show that few-shot learning provides advantages in the in-distribution (ID) scenario, while zero-shot approaches perform better in the out-of-distribution (OOD) scenario. In attack source verification with ID trials, few-shot Siamese and MLP achieve equal error rates (EER) of 17.72% and 13.11%, compared to 29.91% for zero-shot cosine scoring. Conversely, in OOD trials, zero-shot cosine scoring reaches 16.43%, outperforming few-shot Siamese at 23.47% and MLP at 21.57%.

语音伪造零样本溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。