arXiv:2511.09682cs.AIcs.SD2025-11被引 2

提出抗干扰推理训练方法,让语音模型更安全且不降性能。

Rebellion: Noise-Robust Reasoning Training for Audio Reasoning Models

  • 通过对抗性噪声训练,提升模型对复杂攻击的鲁棒性。
  • 在Qwen2-Audio上实现98.7%安全性,同时保持94.2%任务准确率。
  • 适合开发安全语音助手或内容审核系统的研究者参考。

利用推理训练(RT)赋予大型语言模型(LMs)推理能力,显著提升其性能,使音频推理模型(ARMs)日益流行。然而,尚无研究探讨ARMs对旨在诱导有害响应的越狱攻击的安全性。我们首先发现,使用适当的安全推理数据的标准RT可抵御普通越狱攻击,但无法应对我们提出的简单而有效的新型越狱攻击。原因是普通与高级越狱攻击间存在显著表征漂移,迫使目标ARMs输出有害内容。基于此,我们提出Rebellion:一种针对最坏情况表征漂移的鲁棒推理训练方法。所有实验均基于Qwen2-Audio,结果表明Rebellion:1)可在不损害良性任务性能的前提下抵御高级音频越狱攻击;2)显著优于标准RT方法,在准确率-安全性权衡上表现更优。

原文摘要 · Abstract (English)

Instilling reasoning capabilities in large models (LMs) using reasoning training (RT) significantly improves LMs' performances. Thus Audio Reasoning Models (ARMs), i.e., audio LMs that can reason, are becoming increasingly popular. However, no work has studied the safety of ARMs against jailbreak attacks that aim to elicit harmful responses from target models. To this end, first, we show that standard RT with appropriate safety reasoning data can protect ARMs from vanilla audio jailbreaks, but cannot protect them against our proposed simple yet effective jailbreaks. We show that this is because of the significant representation drift between vanilla and advanced jailbreaks which forces the target ARMs to emit harmful responses. Based on this observation, we propose Rebellion, a robust RT that trains ARMs to be robust to the worst-case representation drift. All our results are on Qwen2-Audio; they demonstrate that Rebellion: 1) can protect against advanced audio jailbreaks without compromising performance on benign tasks, and 2) significantly improves accuracy-safety trade-off over standard RT method.

语音模型安全训练越狱攻击推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。