arXiv:2606.10758eess.AS2026-06中稿 · the 34th European …

用代理锚点学习识别语音合成系统,能精准溯源并发现未知模型。

Anchoring the Unknown: Open-Set Model Attribution via Proxy-Anchor Learning

论文配图:Anchoring the Unknown: Open-Set Model Attribution via Proxy-Anchor Learning
图 1 · 摘自论文原文
  • 基于Wav2Vec2-BERT嵌入,用代理锚点损失构建可区分的语音特征空间。
  • 在140个语音系统上达99.76%准确率,未知系统检测误报率仅2.04%。
  • 适合语音取证、反深度伪造研究者,尤其关注开放集场景的应用。

文本到语音(TTS)系统生成逼真合成语音的能力日益增强,给音频鉴伪带来挑战。尽管二分类深度伪造检测已受广泛关注,但源追踪(即识别音频由哪个TTS系统生成)在开放集场景下仍研究不足,尤其当遇到未见系统时。本文提出一种基于代理锚点损失的度量学习框架,作用于Wav2Vec2-BERT嵌入,用于学习具有判别力的嵌入空间,实现TTS源归属与分布外(OOD)检测。我们在涵盖140个TTS系统、51种语言的MLAAD v9数据集上进行评估,并引入一种架构合并策略,将同系统的不同版本归为统一类别,降低类间混淆。系统在110个分布内类别上达到99.76%准确率,分布外检测的假阳性率(FPR@95)低至2.04%。此外,为公平对比当前最优方法,我们在MLAAD v5官方数据集划分上进一步评估,将分布外准确率几乎提升一倍。结果表明,结合架构感知类别设计与后处理分布外评分,代理锚点度量学习为闭集与开集场景下的语音合成源追踪提供了有效方案。

原文摘要 · Abstract (English)

The proliferation of text-to-speech (TTS) systems capable of generating realistic synthetic speech poses growing challenges for audio forensics. While binary deepfake detection has received considerable attention, source tracing (i.e., identifying which TTS system produced a given audio sample) remains underexplored, particularly in open-set scenarios where unknown systems may be encountered. We propose a metric learning framework based on the Proxy-Anchor loss function that operates on Wav2Vec2-BERT embeddings to learn a discriminative embedding space for TTS source attribution and out-of-distribution (OOD) detection of unseen systems. We evaluate it on the MLAAD v9 dataset spanning 140 TTS systems across 51 languages, and introduce an architecture merging strategy that groups TTS system versions into unified classes, reducing inter-class confusion. Our system achieves 99.76% accuracy on 110 in-distribution classes and a False Positive Rate (FPR@95) as low as 2.04% for OOD detection. Also, for a fair comparison against the current state of the art, we further evaluate it on the MLAAD v5 official dataset splits, improving the OOD accuracy by almost doubling it. These results demonstrate that Proxy-Anchor metric learning, combined with architecture-aware class design and post-hoc OOD scoring, provides an effective framework for forensic TTS source tracing in both closed-set and open-set settings.

语音生成源追踪度量学习深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。