arXiv:2508.08559cs.SDcs.LG2025-08中稿 · IEEE Automatic Spe…被引 1

用点击声触发,同时攻击50个说话人识别系统。

Multi-Target Backdoor Attacks Against Speaker Recognition

  • 用不依赖位置的点击声作触发器,可同时攻击多目标。
  • 最高成功率95.04%,在高相似度对下验证攻击达90%。
  • 模拟真实环境噪声,平衡隐蔽性与攻击效果,适合安全研究者。

本文提出一种针对说话人识别的多目标后门攻击方法,使用与位置无关的点击声作为触发信号。不同于以往单目标攻击,该方法可同时针对最多50名说话人,最高成功率可达95.04%。为模拟更真实的攻击场景,我们调整语音与触发信号之间的信噪比,揭示了隐蔽性与有效性之间的权衡。进一步将攻击扩展至说话人验证任务,通过余弦相似度选择最接近的训练说话人作为代理目标。当目标说话人与注册说话人高度相似时,攻击成功率最高可达90%。

原文摘要 · Abstract (English)

In this work, we propose a multi-target backdoor attack against speaker identification using position-independent clicking sounds as triggers. Unlike previous single-target approaches, our method targets up to 50 speakers simultaneously, achieving success rates of up to 95.04%. To simulate more realistic attack conditions, we vary the signal-to-noise ratio between speech and trigger, demonstrating a trade-off between stealth and effectiveness. We further extend the attack to the speaker verification task by selecting the most similar training speaker - based on cosine similarity - as a proxy target. The attack is most effective when target and enrolled speaker pairs are highly similar, reaching success rates of up to 90% in such cases.

后门攻击说话人识别安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。