用音频水印做隐蔽后门,既保音质又提升攻击成功率。
Bloodroot: When Watermarking Turns Poisonous For Stealthy Backdoor
- 将水印设计为触发信号,结合对抗性LoRA微调实现隐蔽攻击。
- 在语音识别与说话人识别任务中,触发成功率超90%,干净样本准确率高。
- 适合关注模型安全与数据版权保护的研究者。
后门数据投毒是所有权保护和防御恶意攻击的关键技术。在训练数据中嵌入隐藏触发器可操控模型输出,实现来源验证并阻止未经授权的使用。然而,现有音频后门方法效果不佳,因污染音频常导致感知质量下降,易被人类听出异常。本文探究了音频水印在实现隐蔽投毒中的内在有效性,提出创新的「水印即触发器」概念,并通过对抗性LoRA微调集成至Bloodroot后门框架中,在显著提升感知质量的同时,实现更高的触发成功率与干净样本准确率。在语音识别(SR)和说话人识别(SID)数据集上的实验表明,基于水印的投毒在声学滤波和模型剪枝下仍保持有效。所提出的Bloodroot框架不仅保障了数据到模型的所有权,也揭示了对抗性滥用的风险。
原文摘要 · Abstract (English)
Backdoor data poisoning is a crucial technique for ownership protection and defending against malicious attacks. Embedding hidden triggers in training data can manipulate model outputs, enabling provenance verification, and deterring unauthorized use. However, current audio backdoor methods are suboptimal, as poisoned audio often exhibits degraded perceptual quality, which is noticeable to human listeners. This work explores the intrinsic stealthiness and effectiveness of audio watermarking in achieving successful poisoning. We propose a novel Watermark-as-Trigger concept, integrated into the Bloodroot backdoor framework via adversarial LoRA fine-tuning, which enhances perceptual quality while achieving a much higher trigger success rate and clean-sample accuracy. Experiments on speech recognition (SR) and speaker identification (SID) datasets show that watermark-based poisoning remains effective under acoustic filtering and model pruning. The proposed Bloodroot backdoor framework not only secures data-to-model ownership, but also well reveals the risk of adversarial misuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。