首个针对音频大模型的越狱攻击评测基准,揭示语音隐秘指令风险。
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
- 构建音频越狱工具箱与多样化攻击数据集,支持隐写注入和编辑。
- 在多个先进音频大模型上验证攻击成功率超60%,暴露严重安全漏洞。
- 适合研究模型安全、语音生成防御的团队使用。
大型语言模型(LLMs)在自然语言处理任务中展现出卓越的零样本性能。通过整合多模态编码器,其能力扩展至视觉与听觉输入,形成多模态大模型(MLLMs)。然而,这些先进能力也可能带来重大安全隐患,因模型可能被越狱攻击诱导生成有害内容。尽管已有研究广泛探索文本与视觉模态的越狱方法,但针对大型音频语言模型(LALMs)的音频特定越狱威胁仍鲜有研究。为此,本文提出Jailbreak-AudioBench,包含工具箱、精选数据集与全面评测基准。工具箱支持文本转音频及多种隐写编辑技术,数据集提供原始与修改后的显性和隐性越狱音频样本。基于此数据集,我们评估了多个主流LALMs,建立了迄今最全面的音频越狱评测基准。该工作为未来LALMs的安全对齐研究奠定基础,可深入揭示如基于查询的音频编辑等新型越狱威胁,并推动有效防御机制的发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate impressive zero-shot performance across a wide range of natural language processing tasks. Integrating various modality encoders further expands their capabilities, giving rise to Multimodal Large Language Models (MLLMs) that process not only text but also visual and auditory modality inputs. However, these advanced capabilities may also pose significant safety problems, as models can be exploited to generate harmful or inappropriate content through jailbreak attacks. While prior work has extensively explored how manipulating textual or visual modality inputs can circumvent safeguards in LLMs and MLLMs, the vulnerability of audio-specific jailbreak on Large Audio-Language Models (LALMs) remains largely underexplored. To address this gap, we introduce Jailbreak-AudioBench, which consists of the Toolbox, curated Dataset, and comprehensive Benchmark. The Toolbox supports not only text-to-audio conversion but also various editing techniques for injecting audio hidden semantics. The curated Dataset provides diverse explicit and implicit jailbreak audio examples in both original and edited forms. Utilizing this dataset, we evaluate multiple state-of-the-art LALMs and establish the most comprehensive Jailbreak benchmark to date for audio modality. Finally, Jailbreak-AudioBench establishes a foundation for advancing future research on LALMs safety alignment by enabling the in-depth exposure of more powerful jailbreak threats, such as query-based audio editing, and by facilitating the development of effective defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。