自动化音频红队测试框架,发现语音大模型隐藏安全漏洞
ARENA: Automated Red-Teaming for Large Audio Language Models

- 构建闭环系统,用音频触发文本安全漏洞的联合攻击
- 在4个模型上最高检测率100%,误报率低于13%
- 适合安全评估、模型审计人员使用
大型语音-语言模型(LALMs)支持通过语音、音乐和环境音与模型交互,但也带来了文本红队难以暴露的安全风险。本文研究基于音频的自动化红队测试,即文本本身安全,但与音频结合后会诱导有害行为。提出ARENA框架,在独立2000例文本-音频数据集上训练控制器,由MD-Judge提供训练奖励与自适应搜索反馈,最终由非自适应的Llama Guard 3单独标注结果。在520个保留的AdvBench目标上,ARENA对Audio Flamingo 3、Qwen2-Audio、MiMo-Audio和GPTAudio的漏报率(FDR)分别为87.9%、71.5%、68.1%、75.4%,精确率(PSR)分别为100.0%、96.3%、100.0%、98.5%。消融实验表明,反馈精炼和音频变体搜索显著提升攻击发现能力。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated audio-grounded red-teaming, where a text query must remain safe in isolation while the joint text-audio input induces harmful target behavior. We propose ARENA, a closed-loop framework that trains a controller on an independent 2,000case text-audio dataset. MD-Judge supplies training rewards and adaptive search feedback, while a separate, non-adaptive Llama Guard 3 evaluator alone labels final outcomes. On 520 held-out AdvBench objectives, ARENA achieves FDR/PSR of 87.9/100.0%, 71.5/96.3%, 68.1/100.0%, and 75.4/98.5% on Audio Flamingo 3, Qwen2-Audio, MiMo-Audio, and GPTAudio, respectively. Ablations show that feedback-based refinement and audio-variant search substantially improve attack discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。