轻量级语音防伪网络,专为物联网设备设计,有效抵御多种语音伪造攻击。
Parallel Stacked Aggregated Network for Voice Authentication in IoT-Enabled Smart Devices
- 分段-变换-聚合架构,直接处理原始音频,无需手工特征
- 在多个伪造攻击上保持稳定性能,等错误率更低
- 适合资源受限的物联网设备部署,通用性强
近年来,随着用户隐私与安全问题日益突出,基于物联网智能设备的语音认证备受关注。现有认证系统易受重放、语音克隆和深度伪造音频等语音欺骗攻击,导致身份冒用、非法访问及金融欺诈等风险。当前解决方案多针对单一攻击类型,对未知攻击泛化能力差;而统一反欺骗方案通常结构复杂,难以在物联网设备上部署,且在特定攻击下表现不佳。为此,本文提出并行堆叠聚合网络(PSA-Net),一种专为语音控制型物联网设备设计的轻量级反欺骗防御框架。PSA-Net 直接处理原始音频,无需依赖数据集的手工特征或预计算频谱图。其采用分段-变换-聚合策略:将语音片段化,通过卷积提取可微分的内在嵌入,并进行聚合以区分真实与伪造语音。相比传统深度残差网络,引入基数(cardinality)作为额外维度,增强模型对多样化攻击的泛化能力。实验表明,PSA-Net 在多种攻击场景下表现出更一致的性能,显著优于现有方法。
原文摘要 · Abstract (English)
Voice authentication on IoT-enabled smart devices has gained prominence in recent years due to increasing concerns over user privacy and security. The current authentication systems are vulnerable to different voice-spoofing attacks (e.g., replay, voice cloning, and audio deepfakes) that mimic legitimate voices to deceive authentication systems and enable fraudulent activities (e.g., impersonation, unauthorized access, financial fraud, etc.). Existing solutions are often designed to tackle a single type of attack, leading to compromised performance against unseen attacks. On the other hand, existing unified voice anti-spoofing solutions, not designed specifically for IoT, possess complex architectures and thus cannot be deployed on IoT-enabled smart devices. Additionally, most of these unified solutions exhibit significant performance issues, including higher equal error rates or lower accuracy for specific attacks. To overcome these issues, we present the parallel stacked aggregation network (PSA-Net), a lightweight framework designed as an anti-spoofing defense system for voice-controlled smart IoT devices. The PSA-Net processes raw audios directly and eliminates the need for dataset-dependent handcrafted features or pre-computed spectrograms. Furthermore, PSA-Net employs a split-transform-aggregate approach, which involves the segmentation of utterances, the extraction of intrinsic differentiable embeddings through convolutions, and the aggregation of them to distinguish legitimate from spoofed audios. In contrast to existing deep Resnet-oriented solutions, we incorporate cardinality as an additional dimension in our network, which enhances the PSA-Net ability to generalize across diverse attacks. The results show that the PSA-Net achieves more consistent performance for different attacks that exist in current anti-spoofing solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。