arXiv:2505.13541eess.AScs.LG2025-05EMNLP被引 6

提出可实时防御语音模型越狱攻击的补丁技术,有效提升安全性。

SPIRIT: Patching Speech Language Models against Jailbreak Attacks

  • 通过修改推理时的模型激活值实现事后防御
  • 在部分攻击下达到99%防御成功率,且不影响使用效果
  • 无需重新训练,适合实际部署场景

语音语言模型(SLMs)通过语音指令实现自然交互,能更精准捕捉用户意图。相比文本模型,语音信号更易被注入人耳难以察觉的噪声以绕过安全机制。我们分析发现,某些情况下SLMs对越狱攻击的防御失败率可达100%。为此,提出一种推理阶段的后处理补丁防御方法,通过干预模型激活值,在不重新训练的前提下,将防御成功率提升至99%,同时保持对模型性能的极小影响。我们在专为语音模型设计的大规模基准上进行了消融实验,验证了该方法的有效性与实用性。

原文摘要 · Abstract (English)

Speech Language Models (SLMs) enable natural interactions via spoken instructions, which more effectively capture user intent by detecting nuances in speech. The richer speech signal introduces new security risks compared to text-based models, as adversaries can better bypass safety mechanisms by injecting imperceptible noise to speech. We analyze adversarial attacks and find that SLMs are substantially more vulnerable to jailbreak attacks, which can achieve a perfect 100% attack success rate in some instances. To improve security, we propose post-hoc patching defenses used to intervene during inference by modifying the SLM's activations that improve robustness up to 99% with (i) negligible impact on utility and (ii) without any re-training. We conduct ablation studies to maximize the efficacy of our defenses and improve the utility/security trade-off, validated with large-scale benchmarks unique to SLMs.

语音模型安全防御越狱攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。