arXiv:2505.24519cs.CVcs.CL2025-05EMNLP被引 7

自动遮蔽图像碎片并分析意图,提升视觉语言模型抗攻击能力

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders

  • 自动屏蔽无关图像区域,干扰恶意扰动
  • 联合分析用户意图,防御成功率从52.4%提至81.7%
  • 无需重训练,兼顾安全与模型通用性

我们提出AMIA,一种轻量级、仅推理的大型视觉语言模型(LVLM)防御机制,具备两项能力:(1) 自动屏蔽少量与文本无关的图像块,以破坏对抗性扰动;(2) 联合执行意图分析,在生成回复前识别并缓解隐藏的有害意图。无需任何微调,AMIA将多种LVLM和越狱测试基准上的防御成功率从平均52.4%提升至81.7%,通用性能仅下降2%,推理开销较小。消融实验表明,遮蔽与意图分析共同构成鲁棒安全-效用平衡的关键。

原文摘要 · Abstract (English)

We introduce AMIA, a lightweight, inference-only defense for Large Vision-Language Models (LVLMs) that (1) Automatically Masks a small set of text-irrelevant image patches to disrupt adversarial perturbations, and (2) conducts joint Intention Analysis to uncover and mitigate hidden harmful intents before response generation. Without any retraining, AMIA improves defense success rates across diverse LVLMs and jailbreak benchmarks from an average of 52.4% to 81.7%, preserves general utility with only a 2% average accuracy drop, and incurs only modest inference overhead. Ablation confirms both masking and intention analysis are essential for a robust safety-utility trade-off.

视觉语言模型安全防御意图分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。