评测大模型在语音攻击下的抗性,发现GPT-4o最抗揍。
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
- 构建四类语音攻击基准,模拟真实对话场景
- 多方法评估显示GPT-4o抗攻击能力最强
- 适合关注语音模型安全的研究者和开发者
对抗性音频攻击对日益普及的大规模音频-语言模型(LALMs)在语音人机交互中的应用构成重大威胁。现有研究多聚焦于特定模型的对抗方法,而实际应用需要更具通用性和普适性的音频对抗攻击方案。本文提出Chat-Audio Attacks(CAA)基准,包含四种不同类型的音频攻击,旨在探究LALMs在对话场景下对这些攻击的脆弱性。为评估LALMs的鲁棒性,我们设计三种评估策略:标准评估,采用传统指标量化模型在攻击下的表现;GPT-4o-based评估,模拟真实对话复杂性;人工评估,揭示用户感知与信任变化。我们在CAA基准上使用三种评估方法,对六种具备语音交互能力的前沿LALMs(包括Gemini-1.5-Pro、GPT-4o等)进行了全面评测。结果表明,四种攻击类型均显著影响模型性能,其中GPT-4o展现出最高抗性。数据可访问:https://github.com/crystraldo/CAA。
原文摘要 · Abstract (English)
Adversarial audio attacks pose a significant threat to the growing use of large audio-language models (LALMs) in voice-based human-machine interactions. While existing research focused on model-specific adversarial methods, real-world applications demand a more generalizable and universal approach to audio adversarial attacks. In this paper, we introduce the Chat-Audio Attacks (CAA) benchmark including four distinct types of audio attacks, which aims to explore the vulnerabilities of LALMs to these audio attacks in conversational scenarios. To evaluate the robustness of LALMs, we propose three evaluation strategies: Standard Evaluation, utilizing traditional metrics to quantify model performance under attacks; GPT-4o-Based Evaluation, which simulates real-world conversational complexities; and Human Evaluation, offering insights into user perception and trust. We evaluate six state-of-the-art LALMs with voice interaction capabilities, including Gemini-1.5-Pro, GPT-4o, and others, using three distinct evaluation methods on the CAA benchmark. Our comprehensive analysis reveals the impact of four types of audio attacks on the performance of these models, demonstrating that GPT-4o exhibits the highest level of resilience. Our data can be accessed via the following link: \href{https://github.com/crystraldo/CAA}{CAA}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。