提出多轮图像输入安全评估框架,发现VLLM在对话中漏洞更深。
REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
- 构建自动化流程,生成对抗性图像并扩展多轮对话
- 多轮测试缺陷率显著高于单轮,最高达16.55%
- 适合关注AI安全、伦理与模型评测的研究者
视觉大语言模型(VLLMs)融合图像处理与文本理解能力,推动AI发展,但其复杂性带来新的安全与伦理挑战,尤其在多模态多轮对话中。传统基于文本的单轮评估方法已不适用。为此,我们提出REVEAL框架,一个可扩展、自动化的图像输入危害评估管道,包含自动图像挖掘、合成对抗数据生成、基于渐进攻击策略的多轮对话扩展,以及通过GPT-4o等评估器进行综合危害检测。我们对五款前沿VLLM(GPT-4o、Llama-3.2、Qwen2-VL、Phi3.5V、Pixtral)在性危害、暴力和误导信息三类风险上进行了评估。结果表明,多轮交互导致缺陷率显著上升;其中Llama-3.2多轮缺陷率达16.55%,Qwen2-VL多轮拒绝率高达19.1%;GPT-4o在安全-可用性指数(SUI)上表现最佳,紧随其后的是Pixtral。误导信息成为亟需加强上下文防御的关键领域。
原文摘要 · Abstract (English)
Vision Large Language Models (VLLMs) represent a significant advancement in artificial intelligence by integrating image-processing capabilities with textual understanding, thereby enhancing user interactions and expanding application domains. However, their increased complexity introduces novel safety and ethical challenges, particularly in multi-modal and multi-turn conversations. Traditional safety evaluation frameworks, designed for text-based, single-turn interactions, are inadequate for addressing these complexities. To bridge this gap, we introduce the REVEAL (Responsible Evaluation of Vision-Enabled AI LLMs) Framework, a scalable and automated pipeline for evaluating image-input harms in VLLMs. REVEAL includes automated image mining, synthetic adversarial data generation, multi-turn conversational expansion using crescendo attack strategies, and comprehensive harm assessment through evaluators like GPT-4o. We extensively evaluated five state-of-the-art VLLMs, GPT-4o, Llama-3.2, Qwen2-VL, Phi3.5V, and Pixtral, across three important harm categories: sexual harm, violence, and misinformation. Our findings reveal that multi-turn interactions result in significantly higher defect rates compared to single-turn evaluations, highlighting deeper vulnerabilities in VLLMs. Notably, GPT-4o demonstrated the most balanced performance as measured by our Safety-Usability Index (SUI) followed closely by Pixtral. Additionally, misinformation emerged as a critical area requiring enhanced contextual defenses. Llama-3.2 exhibited the highest MT defect rate ($16.55 \%$) while Qwen2-VL showed the highest MT refusal rate ($19.1 \%$).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。