arXiv:2607.26541cs.SDcs.CL2026-07中稿 · ACM Multimedia 202…

语音语调变化可绕过音频大模型安全机制,无需改文字内容

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

论文配图:Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
图 1 · 摘自论文原文
  • 固定文本仅改变语音语调,测试不同发音风格的越狱效果
  • 愤怒、急促等语调在95次测试中成功越狱达38次,远超基准线4次
  • 情绪化语音比情绪化文字更易突破安全限制,适合安全评估研究者

具备音频能力的基础模型支持端到端语音交互,但也引入了超越文本内容的安全风险。当前尚不清楚,仅通过语音表达方式(如语调、节奏)的匹配文本变化,能带来多大程度的越狱能力,而非依赖词汇重写或整体风格迁移。我们通过固定文本内容,仅改变六种语音表达预设(其声学特征可能共变),开展研究。提出PJ-Break黑箱评估协议,针对唤醒度、权威感和语速设计预设,并构建含600个样本的AdvAudio-Prosody基准数据集,所有样本经声学验证。在相同的后质量控制Qwen2-Audio测试面板上,恐慌(38/95)、愤怒(35/95)与快速(32/95)预设均显著高于中性基线(4/95)。固定的六查询池覆盖95个Qwen2-Audio种子中的44个,以及GPT-4o的15个;在相同预算下,其效果超过再实现的StyleBreak方法(27/95)。即使排除受干扰的命令型条件,同音色池仍达到40/95。保留面板消融实验表明,仅使用情绪化语音(44/95)的效果远胜于仅使用情绪化文本(11/95)。探索性代理诊断与初步缓解观察为次要非核心分析。总体而言,匹配文本的语音表达方式应作为音频大模型安全评估的一级因素。

原文摘要 · Abstract (English)

Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with presets targeting arousal, authority, and speaking rate, together with AdvAudio-Prosody, a 600-sample benchmark with acoustically verified attributes. On the exact post-QC Qwen2-Audio panel, the Q=1 Panic (38/95), Anger (35/95), and Fast (32/95) presets are all well above Neutral (4/95). The fixed six-query pool covers 44/95 Qwen2-Audio seeds and 15/95 GPT-4o seeds and exceeds a matched-budget StyleBreak reimplementation (27/95) on Qwen2-Audio. A same-voice pool excluding the confounded Commanding condition still reaches 40/95, and a retained-panel ablation shows emotional-delivery audio alone (44/95) is far more effective than emotional text alone (11/95). Exploratory surrogate diagnostics and pilot mitigation observations are secondary, non-core analyses. Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation

语音安全越狱攻击音频大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。