用单帧扰动就能操控多模态大模型的决策,威胁真实场景安全。
Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
- 设计可感知语义的通用扰动,引导模型输出指定目标。
- 在三个模型上实现66%攻击成功率,单帧覆盖五个目标。
- 适合关注AI安全、对抗样本与多模态系统漏洞的研究者。
多模态大语言模型(MLLMs)正广泛部署于自动驾驶、机器人等无状态系统中。本文研究一种新型威胁:语义感知劫持。我们探索利用单一通用扰动同时劫持多个无状态决策的可行性。提出语义感知通用扰动(SAUP),作为语义路由器,主动感知输入语义并将其导向攻击者定义的目标。基于对潜在空间几何特性的理论与实证分析,设计语义导向优化策略(SORT),并构建包含细粒度语义标注的新数据集以评估性能。在三个代表性MLLM上的大量实验表明,该攻击具备根本可行性,在仅使用单帧的情况下,对Qwen模型实现了66%的攻击成功率,覆盖五个目标。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are increasingly deployed in stateless systems, such as autonomous driving and robotics. This paper investigates a novel threat: Semantic-Aware Hijacking. We explore the feasibility of hijacking multiple stateless decisions simultaneously using a single universal perturbation. We introduce the Semantic-Aware Universal Perturbation (SAUP), which acts as a semantic router, "actively" perceiving input semantics and routing them to distinct, attacker-defined targets. To achieve this, we conduct theoretical and empirical analysis on the geometric properties in the latent space. Guided by these insights, we propose the Semantic-Oriented (SORT) optimization strategy and annotate a new dataset with fine-grained semantics to evaluate performance. Extensive experiments on three representative MLLMs demonstrate the fundamental feasibility of this attack, achieving a 66% attack success rate over five targets using a single frame against Qwen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。