用音色泄露技术隐蔽植入语音分类模型后门,一次训练多靶点攻击。
Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

- 通过元学习与冲突梯度投影,一次性注入多个隐蔽后门。
- 攻击成功率超基线方法,且样本听感自然、难被检测。
- 适合研究模型安全或防御机制的人员关注。
近年来,语音分类方法在智能设备中广泛应用。现有研究表明,后门攻击对这些模型构成严重安全威胁,亟需探索新型攻击手段以暴露和防范风险。本文分析了当前语音触发器易被深度神经网络防御者检测的问题,提出音色泄露攻击(Timbre Leakage Attack, TLA)。该触发器在深层自监督特征的帧级别传播音色信息,生成人类感知上自然的污染样本。进一步提出Pmeta-TLA,一种基于元学习与投影冲突梯度(PCGrad)的多后门注入训练机制,将TLA作为多目标攻击工具。在关键词识别任务中,针对多种深度神经网络模型进行数据投毒攻击测试。实验表明,该策略在攻击效果、隐蔽性、鲁棒性及攻击成本方面均优于基线方法。
原文摘要 · Abstract (English)
Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attacks provide a substantial security concern to these models, underscoring the pressing necessity to investigate additional potential attack techniques to expose and prevent such risks. This work discusses the vulnerability of current speech triggers to detection by deep neural network defenders and introduces the Timbre Leakage Attack (TLA). The suggested trigger disseminates timbre information at the frame level within the deep self-supervised features, producing poisoned samples that appear natural to human perception. Furthermore, we introduce Pmeta-TLA, an innovative training mechanism for embedding numerous backdoors one time. This method proposes a multi-backdoor injection training strategy using meta-learning and Projected Conflicting Gradients (PCGrad) and introduces TLA as a multi-target attack tool within it. We performed tests on data-poisoning backdoor attacks in keyword spotting tasks utilizing some deep neural network models. Experimental results indicate that the proposed strategy attains superior Attack efficacy, enhanced stealthiness, robustness, and a reduced attack cost relative to baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。