arXiv:2606.07494cs.SDeess.AS2026-06

通过随机化特征统计模拟真实场景,提升伪造语音检测泛化能力。

Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech

论文配图:Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech
图 1 · 摘自论文原文
  • 训练时将固定特征转为随机分布,模拟真实语音多样性
  • 在40种未知生成模型上实现领先检测效果
  • 适合需要强泛化能力的语音安全防护研究者

基于神经音频编解码器的语音生成(CodecFake)生成高度逼真的音频,对现有伪造语音检测模型构成挑战。尽管使用编解码重合成语音(CoRS)作为代理数据可提升性能,但泛化能力有限。本文提出领域偏移特征增强(DSFA),在微调阶段将确定性特征统计转化为随机分布,以模拟真实场景中的变化。为评估泛化能力,我们进一步引入编解码语音生成扩展评估(CoSG ExtEval)数据集,该数据集是CoSG Eval(来自CodecFake+)的更难版本,包含40种未见生成模型和长语音片段。实验表明,结合后训练的自监督学习骨干网络与DSFA,能有效缩小代理数据与真实场景间的领域差距。该方法在CoSG Eval和CoSG ExtEval上的多种CodecFake攻击中均达到当前最优性能。

原文摘要 · Abstract (English)

Recent neural audio codec-based speech generation (CodecFake) produces highly realistic audio, posing a challenge to existing deepfake countermeasure models. While using codec resynthesized speech (CoRS) as proxy data improves performance, it often suffers from limited generalization. We propose Domain-Shift Feature Augmentation (DSFA), which simulates "in-the-wild" variations by transforming deterministic feature statistics into stochastic distributions during fine-tuning. To evaluate generalization, we further introduce Codec-based Speech Generation Extension Evaluation (CoSG ExtEval) dataset, a more challenging extension of the CoSG Eval (from CodecFake+) dataset, featuring 40 unseen generative models and long-form audio. Experimental results demonstrate that combining a post-trained SSL backbone with DSFA effectively narrows the proxy-to-wild domain gap. This approach achieves state-of-the-art performance across diverse CodecFake attacks in both CoSG Eval and CoSG ExtEval.

语音伪造深度伪造域泛化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。