多模态排版攻击让音视频大模型易被欺骗,成功率超80%。
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
- 设计跨模态文字干扰,同时影响音视频与文本输入
- 协同攻击使成功率提升至83.43%,远高于单一模态攻击
- 适合安全研究者、大模型防御团队参考
随着音视频多模态大语言模型(MLLMs)在安全关键场景中日益普及,理解其漏洞至关重要。本文提出多模态排版攻击,系统研究跨模态干扰对MLLMs的影响。不同于以往聚焦单模态攻击的研究,我们揭示了多模态模型在音频、视觉与文本扰动协同下的脆弱性。实验表明,联合多模态攻击的攻击成功率达83.43%,显著高于单模态攻击的34.93%。该发现覆盖多个前沿MLLMs、任务及常识推理与内容审核基准,确立多模态排版攻击为多模态推理中关键且被忽视的威胁。代码与数据将公开共享。
原文摘要 · Abstract (English)
As audio-visual multi-modal large language models (MLLMs) are increasingly deployed in safety-critical applications, understanding their vulnerabilities is crucial. To this end, we introduce Multi-Modal Typography, a systematic study examining how typographic attacks across multiple modalities adversely influence MLLMs. While prior work focuses narrowly on unimodal attacks, we expose the cross-modal fragility of MLLMs. We analyze the interactions between audio, visual, and text perturbations and reveal that coordinated multi-modal attack creates a significantly more potent threat than single-modality attacks (attack success rate = $83.43\%$ vs $34.93\%$).Our findings across multiple frontier MLLMs, tasks, and common-sense reasoning and content moderation benchmarks establishes multi-modal typography as a critical and underexplored attack strategy in multi-modal reasoning. Code and data will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。