arXiv:2506.11521cs.CRcs.AI2025-06综述被引 4

系统梳理多模态模型音频视觉攻击与防御,揭示安全漏洞

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models

  • 系统分类音频视觉攻击类型,涵盖对抗、后门和越狱攻击
  • 指出当前多模态模型易被指令操控生成有害内容
  • 适合关注AI安全、多模态模型研究者阅读

多模态大语言模型(MLLMs)在音视频与自然语言处理间架起桥梁,在多个音视频任务中表现优异。然而,高质量音视频训练数据和算力资源稀缺,导致研究者广泛依赖第三方数据和开源模型,埋下安全隐患。实证研究表明,最新MLLMs可被指令或输入(如对抗扰动、恶意查询)操控,生成有害内容,且能绕过模型内部安全机制。为深入理解音视频多模态模型的安全风险,本文系统综述了对抗攻击、后门攻击和越狱攻击等各类攻击形式。现有综述多局限于特定攻击类型,缺乏统一框架。本文首次全面覆盖最新音视频多模态模型中的多种攻击,提出未来研究挑战与趋势,为安全防御提供关键参考。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs), which bridge the gap between audio-visual and natural language processing, achieve state-of-the-art performance on several audio-visual tasks. Despite the superior performance of MLLMs, the scarcity of high-quality audio-visual training data and computational resources necessitates the utilization of third-party data and open-source MLLMs, a trend that is increasingly observed in contemporary research. This prosperity masks significant security risks. Empirical studies demonstrate that the latest MLLMs can be manipulated to produce malicious or harmful content. This manipulation is facilitated exclusively through instructions or inputs, including adversarial perturbations and malevolent queries, effectively bypassing the internal security mechanisms embedded within the models. To gain a deeper comprehension of the inherent security vulnerabilities associated with audio-visual-based multimodal models, a series of surveys investigates various types of attacks, including adversarial and backdoor attacks. While existing surveys on audio-visual attacks provide a comprehensive overview, they are limited to specific types of attacks, which lack a unified review of various types of attacks. To address this issue and gain insights into the latest trends in the field, this paper presents a comprehensive and systematic review of audio-visual attacks, which include adversarial attacks, backdoor attacks, and jailbreak attacks. Furthermore, this paper also reviews various types of attacks in the latest audio-visual-based MLLMs, a dimension notably absent in existing surveys. Drawing upon comprehensive insights from a substantial review, this paper delineates both challenges and emergent trends for future research on audio-visual attacks and defense.

多模态安全对抗攻击模型防御音视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。