为语音反伪造模型提供概率化鲁棒性验证,防未知合成攻击
Probabilistic Verification of Voice Anti-Spoofing Models
- 基于概率方法评估语音伪造检测模型在TTS、克隆等攻击下的误判率
- 理论推导出误分类概率上界,在多种攻击下验证有效
- 不依赖具体模型,适合评估未知生成技术的防御能力
生成模型的进展加剧了语音合成技术被恶意滥用的风险,攻击者可模仿目标说话人访问敏感资源。尽管语音深度伪造检测已快速进步,但多数现有对策缺乏形式化鲁棒性保障,或无法泛化到未见的生成技术。本文提出PV-VASM,一种针对语音反伪造模型(VASMs)的概率化鲁棒性验证框架。该方法估计在文本转语音(TTS)、语音克隆(VC)及参数化信号变换下的误分类概率。该方法具有模型无关性,可验证对未见语音合成技术及输入扰动的鲁棒性。我们推导了错误概率的理论上限,并在多种实验设置中验证了该方法的有效性,证明其作为实用鲁棒性验证工具的价值。
原文摘要 · Abstract (English)
Recent advances in generative models have amplified the risk of malicious misuse of speech synthesis technologies, enabling adversaries to impersonate target speakers and access sensitive resources. Although speech deepfake detection has progressed rapidly, most existing countermeasures lack formal robustness guarantees or fail to generalize to unseen generation techniques. We propose PV-VASM, a probabilistic framework for verifying the robustness of voice anti-spoofing models (VASMs). PV-VASM estimates the probability of misclassification under text-to-speech (TTS), voice cloning (VC), and parametric signal transformations. The approach is model-agnostic and enables robustness verification against unseen speech synthesis techniques and input perturbations. We derive a theoretical upper bound on the error probability and validate the method across diverse experimental settings, demonstrating its effectiveness as a practical robustness verification tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。