通过信号级融合实现语音身份伪造,严重威胁语音验证系统安全
Time-Domain Voice Identity Morphing (TD-VIM): A Signal-Level Approach to Morphing Attacks on Speaker Verification Systems
- 在时域直接混合双人语音特征,生成可冒充多身份的伪造音频
- 在文本依赖场景下,攻击成功率高达99.74%(三星S8)
- 适用于研究语音安全漏洞或评估生物识别系统抗攻击能力的人员
在生物识别系统中,通常将每个样本与特定个体关联。然而,近期研究已证明可生成能匹配多个身份的‘融合’生物特征样本,此类融合攻击被视为生物识别系统的潜在安全风险。现有研究多集中于图像模态(如人脸、指纹、虹膜),而本文提出时域语音身份融合(TD-VIM)方法,首次在信号层面实现两段不同语音身份特征的融合,生成对语音验证系统具有高欺骗性的伪造样本。基于多语言音视频智能手机数据库,本研究构建了四类融合语音信号,并通过全面脆弱性分析评估其效果。采用通用融合攻击潜力(G-MAP)指标,在两个基于深度学习的语音验证系统及一个商用系统Verispeak上进行测试。结果表明,融合语音样本在文本依赖场景下,于0.1%误匹配率条件下,苹果iPhone-11和三星S8设备上的攻击成功率分别达99.40%和99.74%。
原文摘要 · Abstract (English)
In biometric systems, it is a common practice to associate each sample or template with a specific individual. Nevertheless, recent studies have demonstrated the feasibility of generating "morphed" biometric samples capable of matching multiple identities. These morph attacks have been recognized as potential security risks for biometric systems. However, most research on morph attacks has focused on biometric modalities that operate within the image domain, such as the face, fingerprints, and iris. In this work, we introduce Time-domain Voice Identity Morphing (TD-VIM), a novel approach for voice-based biometric morphing. This method enables the blending of voice characteristics from two distinct identities at the signal level, creating morphed samples that present a high vulnerability for speaker verification systems. Leveraging the Multilingual Audio-Visual Smartphone database, our study created four distinct morphed signals based on morphing factors and evaluated their effectiveness using a comprehensive vulnerability analysis. To assess the security impact of TD-VIM, we benchmarked our approach using the Generalized Morphing Attack Potential (G-MAP) metric, measuring attack success across two deep-learning-based Speaker Verification Systems (SVS) and one commercial system, Verispeak. Our findings indicate that the morphed voice samples achieved a high attack success rate, with G-MAP values reaching 99.40% on iPhone-11 and 99.74% on Samsung S8 in text-dependent scenarios, at a false match rate of 0.1%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。