提出自适应框架,同时攻击语音大模型的文本与音频漏洞。
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs
- 用反馈引导的变异引擎,自动生成文本和音频扰动
- 在6个系统上测试,均发现显著安全漏洞
- 适合研究语音模型安全或对抗攻防的开发者
随着大语言模型(LLMs)越来越多地集成到语音应用中,其面临语音对抗攻击的风险日益突出。现有系统主要分为两类:级联流水线(先将语音转为文本再处理)和端到端大音频-语言模型(LALMs,直接解析原始音频)。前者易受语音传递的文本级越狱攻击,后者还存在声学-语义层面的额外攻击路径。但现有研究多局限于单一架构,对整体语音攻击空间覆盖不足。为此,我们提出一个统一实验环境下,适用于级联流水线与LALMs的自适应越狱攻击框架。核心是反馈引导的变异引擎,可自动生成并优化文本提示与音频扰动,提升攻击多样性与覆盖率。在六个代表性语音系统上的实验表明,两类架构均严重易受音频越狱攻击。相比现有最优方法,本框架在多种语音大模型上实现更一致的更高成功率。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly integrated into audio-based applications, growing concerns have emerged regarding their vulnerability to audio-based adversarial attacks. These systems typically follow two architectural paradigms: cascaded pipelines, where automatic speech recognition converts audio inputs into text before LLM processing, and end-to-end large audio-language models (LALMs), which directly interpret raw audio signals. Beyond architectural differences, cascaded pipelines are primarily vulnerable to text-level jailbreak strategies delivered through speech, whereas end-to-end LALMs introduce additional acoustic-semantic attack vectors. However, existing studies often focus on a single paradigm and provide limited coverage of the broader audio attack space. To bridge this gap, we propose an adaptive jailbreak attack framework for systematic evaluation of both cascaded pipelines and LALMs under a unified experimental setting. At its core, the framework uses a feedback-guided mutation engine to automatically generate and refine jailbreak candidates across both textual prompts and audio perturbations, thereby expanding attack diversity and coverage. Experiments on six representative audio-based systems demonstrate that both paradigms remain substantially vulnerable to audio jailbreak attacks. Compared with state-of-the-art methods, our framework achieves consistently higher attack success rates across diverse audio-based LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。