arXiv:2606.10223cs.SDcs.AI2026-06被引 2

提出双分支门控融合模型,提升语音伪造溯源的准确性与泛化能力。

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

论文配图:Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing
图 1 · 摘自论文原文
  • 用XLSR-53和66维多维度特征CORES并行提取语音合成痕迹。
  • 在MLAAD数据集上实现97.6%识别准确率,误报率降低83.5%。
  • 门控机制自适应融合特征,适合应对未知合成器的开放集场景。

将合成语音追溯到其生成系统仍是未解难题:封闭集模型无法拒绝未知合成器,且预测过于自信。为此,我们提出双分支门控融合框架,结合XLSR-53与CORES——一种66维描述符,涵盖倒谱、振荡、节奏、能量与频谱等维度,可捕捉互补的合成痕迹。分析表明,XLSR-53在域内(ID)保持判别力,而CORES在分布偏移(OOD)下稳定泛化,但简单拼接因自监督表示不平衡而失效。为此,设计输入感知门控,在联合训练中结合交叉熵、能量边界损失(用于区分ID/OOD)与门控多样性项。在MLAAD基准上,系统达到97.6%的ID准确率,4.9%的EERc,相比Interspeech 2025基线,相对FPR95降低83.5%。

原文摘要 · Abstract (English)

Attributing a synthetic utterance to its originating system remains an open challenge: closed-set models fail to reject unseen synthesizers and produce overconfident predictions. To address this, we propose a dual-branch gated fusion framework that pairs XLSR-53 with CORES, a 66-dimensional descriptor that, unlike prior Linear Filter Bank (LFB)-only work, spans cepstral, oscillatory, rhythmic, energy, and spectral dimensions to capture complementary synthesis artifacts. Our analysis shows XLSR-53 remains discriminative in-domain (ID) while CORES generalizes stably under distribution shift (OOD), yet their naive concatenation fails due to SSL representational imbalance. To resolve this, an input-conditioned gate adaptively weights each branch under joint training with cross-entropy, an energy margin loss for ID/OOD separation, and a gate diversity term. On the MLAAD benchmark, our system achieves 97.6\% ID accuracy, 4.9\% EERc, and an 83.5\% relative FPR95 reduction over the Interspeech 2025 baseline.

语音伪造溯源深度学习开放集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。