音频干扰可让三模态模型失效,最高攻击成功率96%
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
- 仅用音频扰动攻击多模态模型不同处理阶段
- 在低感知失真下实现高达96%的攻击成功率
- 适合关注多模态安全与防御的研究者
集成音频、视觉和语言的多模态基础模型在推理与生成任务中表现强劲,但其对抗性攻击的鲁棒性仍不明确。本文研究一种现实且未被充分探索的威胁模型:针对三模态音视频语言模型的无目标、纯音频对抗攻击。分析了六种互补的攻击目标,覆盖音频编码器表征、跨模态注意力、隐藏状态及输出似然等多阶段。在三个前沿模型和多个基准上,发现纯音频扰动可引发严重多模态失效,攻击成功率最高达96%。攻击可在低感知失真(LPIPS ≤ 0.08,SI-SNR ≥ 0)下成功,且优化时长比数据量提升更有效。跨模型与编码器的迁移性有限;而语音识别系统如Whisper主要响应扰动幅度,在严重失真下攻击成功率超97%。结果揭示多模态系统中被忽视的单模态攻击面,推动需强化跨模态一致性的防御机制。
原文摘要 · Abstract (English)
Multimodal foundation models that integrate audio, vision, and language achieve strong performance on reasoning and generation tasks, yet their robustness to adversarial manipulation remains poorly understood. We study a realistic and underexplored threat model: untargeted, audio-only adversarial attacks on trimodal audio-video-language models. We analyze six complementary attack objectives that target different stages of multimodal processing, including audio encoder representations, cross-modal attention, hidden states, and output likelihoods. Across three state-of-the-art models and multiple benchmarks, we show that audio-only perturbations can induce severe multimodal failures, achieving up to 96% attack success rate. We further show that attacks can be successful at low perceptual distortions (LPIPS <= 0.08, SI-SNR >= 0) and benefit more from extended optimization than increased data scale. Transferability across models and encoders remains limited, while speech recognition systems such as Whisper primarily respond to perturbation magnitude, achieving >97% attack success under severe distortion. These results expose a previously overlooked single-modality attack surface in multimodal systems and motivate defenses that enforce cross-modal consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。