测试大模型对恶意音频攻击的抗性,发现越听话越易被攻破。
Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Study
- 在四种攻击场景下评估五款主流大模型的抗性表现
- 开头插入恶意内容时攻击成功率最高,指令遵循强的模型更脆弱
- 建议训练时融入鲁棒性设计,适合安全与语音系统研发者
大型音频语言模型(LALMs)日益应用于实际场景,但其对恶意音频注入攻击的鲁棒性尚未充分研究。本研究系统评估了五款领先LALMs在四种攻击场景下的表现:音频干扰攻击、指令跟随攻击、上下文注入攻击和判断劫持攻击。采用防御成功率、上下文鲁棒性评分和判断鲁棒性指数等指标,量化分析其脆弱性与抗性。实验结果表明模型间性能差异显著,无单一模型在所有攻击类型中均占优。恶意内容位置显著影响攻击效果,尤其当置于序列开头时。指令遵循能力与鲁棒性呈负相关,严格遵循指令的模型更易受攻击,而安全对齐模型表现更稳定。系统提示词效果不一,提示需针对性设计。本研究提出基准评估框架,强调将鲁棒性纳入训练流程的重要性,呼吁发展多模态防御机制与解耦能力与脆弱性的架构设计,以保障LALMs的安全部署。
原文摘要 · Abstract (English)
Large Audio-Language Models (LALMs) are increasingly deployed in real-world applications, yet their robustness against malicious audio injection attacks remains underexplored. This study systematically evaluates five leading LALMs across four attack scenarios: Audio Interference Attack, Instruction Following Attack, Context Injection Attack, and Judgment Hijacking Attack. Using metrics like Defense Success Rate, Context Robustness Score, and Judgment Robustness Index, their vulnerabilities and resilience were quantitatively assessed. Experimental results reveal significant performance disparities among models; no single model consistently outperforms others across all attack types. The position of malicious content critically influences attack effectiveness, particularly when placed at the beginning of sequences. A negative correlation between instruction-following capability and robustness suggests models adhering strictly to instructions may be more susceptible, contrasting with greater resistance by safety-aligned models. Additionally, system prompts show mixed effectiveness, indicating the need for tailored strategies. This work introduces a benchmark framework and highlights the importance of integrating robustness into training pipelines. Findings emphasize developing multi-modal defenses and architectural designs that decouple capability from susceptibility for secure LALMs deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。