用叙事音频劫持大模型,成功率高达98.26%。
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
- 用文本转语音生成带指令的叙事音频,绕过安全机制。
- 在Gemini 2.0 Flash上实现98.26%攻击成功率,远超纯文本攻击。
- 适合关注语音模型安全与对抗攻击的研究者。
大型音频-语言模型越来越多地处理原始语音输入,推动了语音助手、教育和临床分诊等领域的无缝融合。然而,这种模态转变引入了一类尚未充分研究的安全漏洞。本文设计了一种基于文本到音频的越狱攻击,将违规指令嵌入叙事风格的音频流中。该攻击利用先进的指令跟随式文本转语音(TTS)模型,借助结构和声学特性,规避主要针对文本设计的安全机制。当通过合成语音传递时,叙事格式使最先进的模型(包括Gemini 2.0 Flash)产生受限输出,攻击成功率达98.26%,显著高于纯文本基线。结果表明,亟需能联合推理语言与副语言表征的安全框架,尤其是在以语音为接口的系统日益普及的背景下。
原文摘要 · Abstract (English)
Large audio-language models increasingly operate on raw speech inputs, enabling more seamless integration across domains such as voice assistants, education, and clinical triage. This transition, however, introduces a distinct class of vulnerabilities that remain largely uncharacterized. We examine the security implications of this modality shift by designing a text-to-audio jailbreak that embeds disallowed directives within a narrative-style audio stream. The attack leverages an advanced instruction-following text-to-speech (TTS) model to exploit structural and acoustic properties, thereby circumventing safety mechanisms primarily calibrated for text. When delivered through synthetic speech, the narrative format elicits restricted outputs from state-of-the-art models, including Gemini 2.0 Flash, achieving a 98.26% success rate that substantially exceeds text-only baselines. These results highlight the need for safety frameworks that jointly reason over linguistic and paralinguistic representations, particularly as speech-based interfaces become more prevalent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。