arXiv:2607.01729cs.AIcs.SD2026-07

用强化学习设计隐蔽语音后门攻击,不改标签也能骗过模型

DRL-CLBA: A Clean Label Backdoor Attack for Speech Classification via DDPG Reinforcement Learning

论文配图:DRL-CLBA: A Clean Label Backdoor Attack for Speech Classification via DDPG Reinforcement Learning
图 1 · 摘自论文原文
  • 用DDPG强化学习优化音频触发器位置,实现精准污染
  • 在3个数据集上攻击成功率超90%,且可绕过多种防御
  • 适合研究语音安全的人员,揭示智能语音系统漏洞

语音分类深度学习模型易受后门攻击,恶意触发器会在推理时导致误分类。现有样本特定攻击多依赖污染标签,易被人工数据检测发现。本文提出DRL-CLBA,一种基于深度确定性策略梯度(DDPG)强化学习的清洁标签后门攻击方法。通过深度音频隐写技术将样本特定触发器嵌入原始音频,生成模型深层特征空间中的锚点。强化学习框架能有效将目标样本优化至携带触发器的锚点位置,实现无需标签迁移的污染。在三个数据集和四种不同DNN上的实验表明,该攻击成功率高,可有效绕过部分后门防御。攻击对微调、剪枝及频谱特征防御均表现出强鲁棒性,暴露出语音控制系统中的关键安全隐患。

原文摘要 · Abstract (English)

Deep learning models for speech classification are vulnerable to backdoor attacks, where malicious triggers cause misclassification at inference time. While sample-specific attacks can bypass many defenses, they often rely on poisoned label attack, making them detectable via manual data defense. In this paper, we propose DRL-CLBA, a novel clean label backdoor attack for speech classification that leverages Deep Deterministic Policy Gradient (DDPG) reinforcement learning. We also utilize deep audio steganography to embed sample-specific triggers into source audio, creating feature-space anchors. The proposed reinforcement learning framework effectively optimizes target samples toward trigger-bearing anchor points in the model's deep latent space, enabling label-migration-free poisoning of target samples. Experimental results across three datasets and four different DNNs demonstrate that DRL-CLBA achieves a high attack success rate, effectively bypassing some backdoor defenses. The attack demonstrates strong resistance against fine-tuning, pruning, and spectral signature defenses, exposing critical vulnerabilities in speech-controlled systems.

后门攻击语音安全强化学习深度隐写

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。