融合声学物理特征与不确定性感知,提升边缘语音认证抗深度伪造能力
Audio Physical Dynamics Inspired Deepfake Detection for Voice Authentication Systems
- 结合声带动力学物理特征与自监督学习表征
- 通过贝叶斯集成实现音频不确定性估计
- 适用于对抗边缘侧深度伪造与控制面投毒攻击
部署于网络边缘的语音认证系统面临双重威胁:一是复杂的深度伪造合成攻击,二是分布式联邦学习协议中的控制平面投毒。本文提出一种框架,将基于声学物理动力学的深度伪造检测与边缘学习中的不确定性感知相结合。该框架融合可解释的物理特征(建模声道动力学)与自监督学习模块提取的表征,经简化版多层感知机处理后,通过贝叶斯集成生成不确定性估计。结合音频物理特性评估与样本不确定性,所提框架在抵御先进深度伪造攻击方面保持鲁棒性;同时,基于信任的聚合协议有效防范了边缘语音认证系统中的控制平面投毒。
原文摘要 · Abstract (English)
Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-plane poisoning in distributed federated learning protocols. We present a framework coupling audio physical dynamics deepfake detection with uncertainty-aware in edge learning. The framework fuses interpretable physics features modeling vocal tract dynamics with representations coming from a self-supervised learning module. The representations are then processed via a streamlined Multi-Layer Perceptron backbone, followed by a Bayesian ensemble providing uncertainty estimates. Incorporating audio physical characteristics evaluations and uncertainty estimates of audio samples allows our proposed framework to remain robust to advanced deepfake attacks, while our trust-based aggregation protocol secures the control plane against poisoning in network edge voice authentication systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。