arXiv:2601.02914cs.SDcs.CR2026-01被引 1

语音克隆小样本即可骗过主流声纹验证系统

Vulnerabilities of Audio-Based Biometric Authentication Systems Against Deepfake Speech Synthesis

  • 用极少量语音样本训练的深度伪造模型可绕过商业声纹验证
  • 反欺骗检测器跨合成方法泛化能力差,真实场景鲁棒性显著下降
  • 适合关注生物识别安全、语音伪造防御的研究者与工程师

随着语音深度伪造从研究工具演变为广泛可用的商用产品,高安全等级行业面临严峻的生物识别认证威胁。本文基于大规模语音合成数据集,对当前最先进的说话人认证系统进行了系统性实证评估,揭示两大安全漏洞:1)仅需极少量样本训练的现代语音克隆模型即可轻易绕过商用说话人验证系统;2)反欺骗检测器在不同语音合成方法间泛化能力不足,导致其域内性能与真实场景鲁棒性之间存在显著差距。这些发现呼吁重新审视安全策略,强调需推动架构创新、自适应防御机制,并向多因素认证过渡。

原文摘要 · Abstract (English)

As audio deepfakes transition from research artifacts to widely available commercial tools, robust biometric authentication faces pressing security threats in high-stakes industries. This paper presents a systematic empirical evaluation of state-of-the-art speaker authentication systems based on a large-scale speech synthesis dataset, revealing two major security vulnerabilities: 1) modern voice cloning models trained on very small samples can easily bypass commercial speaker verification systems; and 2) anti-spoofing detectors struggle to generalize across different methods of audio synthesis, leading to a significant gap between in-domain performance and real-world robustness. These findings call for a reconsideration of security measures and stress the need for architectural innovations, adaptive defenses, and the transition towards multi-factor authentication.

语音伪造声纹识别安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。