用进化算法生成能绕过语音大模型安全限制的自然干扰音频
ERIS: Evolutionary Real-world Interference Scheme for Jailbreaking Audio Large Models
- 通过遗传算法优化自然声景中的恶意指令融合
- 在多个语音大模型上实现超越文本和音频基线的越狱效果
- 揭示真实环境噪声可被利用于安全攻击,适合安全研究者参考
现有语音大模型对齐主要关注干净输入,忽视复杂环境下的安全风险。本文提出ERIS框架,将真实世界干扰转化为策略性优化的越狱载体。不同于依赖人工设计声学模式的方法,ERIS采用遗传算法优化自然信号的选择与合成,通过种群初始化、交叉融合和概率变异,演化出融合恶意指令的音频。对人类和安全过滤器而言,这些样本表现为带有无害背景噪声的自然语音,却能绕过对齐机制。在多个语音大模型上的评估显示,ERIS显著优于文本与音频越狱基线。研究揭示看似无害的真实世界干扰可被用于规避安全约束,为复杂声学场景下的防御机制提供新思路。
原文摘要 · Abstract (English)
Existing Audio Large Models (ALMs) alignment focuses on clean inputs, neglecting security risks in complex environments. We propose ERIS, a framework transforming real-world interference into a strategically optimized carrier for jailbreaking ALMs. Unlike methods relying on manually designed acoustic patterns, ERIS uses a genetic algorithm to optimize the selection and synthesis of naturalistic signals. Through population initialization, crossover fusion, and probabilistic mutation, it evolves audio fusing malicious instructions with real-world interference. To humans and safety filters, these samples present as natural speech with harmless background noise, yet bypass alignment. Evaluations on multiple ALMs show ERIS significantly outperforms both text and audio jailbreak baselines. Our findings reveal that seemingly innocuous real-world interference can be leveraged to circumvent safety constraints, providing new insights for defensive mechanisms in complex acoustic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。