arXiv:2606.16532cs.SDcs.AI2026-06中稿 · Interspeech 2026, …

通过双粒度正交解耦提升语音伪造检测泛化能力

Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection

论文配图:Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection
图 1 · 摘自论文原文
  • 在样本和批次两级强制特征正交,消除身份泄漏
  • 跨数据集测试中错误率降低2.60个百分点
  • 无需额外网络或对抗训练,适合实际部署

语音深度伪造检测器常因学习说话人身份特征而非合成伪影而泛化能力差,即隐式身份泄露。现有方法虽能缓解但带来架构复杂或训练不稳问题。本文提出双粒度正交解耦框架,在样本级通过余弦正交性实现方向去相关,在批次级通过交叉协方差正则化消除嵌入维度间的线性相关。采用课程解耦调度逐步强化正交约束,无需辅助网络或对抗机制。在ASVspoof 2019 LA、ASVspoof 2021 DF和In-the-Wild数据集上,分别达到1.35%、7.88%和21.58%的等错误率(EER),在跨数据集迁移任务中比梯度反向解耦提升2.60%绝对误差。

原文摘要 · Abstract (English)

Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage. Existing methods address this but incur architectural complexity or training instability. This paper proposes a dual-granularity orthogonal disentanglement framework enforcing feature independence at two levels: sample-level cosine orthogonality captures directional decorrelation, while batch-level cross-covariance regularization eliminates linear correlations across embedding dimensions. A curriculum disentanglement schedule progressively strengthens the orthogonality constraint without auxiliary networks or adversarial dynamics. Experiments on ASVspoof 2019 LA, ASVspoof 2021 DF, and In-the-Wild datasets demonstrate that the proposed method achieves 1.35%, 7.88%, and 21.58% equal error rates (EER), respectively, surpassing gradient reversal disentanglement by 2.60% absolute on cross-dataset transfer.

语音伪造检测正交解耦泛化能力深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。