arXiv:2502.07905cs.CVcs.LG2025-02被引 3

攻击深求模型引发精准幻觉,揭示视觉嵌入漏洞。

DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities

  • 通过优化图像嵌入诱导目标幻觉,实现高精度控制。
  • 在开放问题上幻觉率高达98.0%,保持图像质量(SSIM>0.88)。
  • 提出多提示检测框架,适合关注模型安全的研究者。

多模态大语言模型(MLLMs)代表了人工智能的前沿技术,其中深求模型作为领先的开源替代方案,性能可媲美闭源系统。尽管这些模型表现卓越,其视觉-语言融合机制却引入特定漏洞。我们对DeepSeek Janus实施改进的嵌入操控攻击,通过系统优化图像嵌入,诱导出目标视觉幻觉。在COCO、DALL-E 3和SVIT数据集上的大量实验表明,在开放式问题中幻觉率最高达98.0%,同时保持高图像保真度(SSIM > 0.88)。分析显示,1B和7B两个版本的DeepSeek Janus均易受此类攻击,闭式评估的幻觉率高于开放式提问。我们引入基于LLaMA-3.1 8B Instruct的新型多提示幻觉检测框架,用于鲁棒评估。鉴于深求模型的开源特性与广泛部署潜力,该研究凸显了在嵌入层加强安全防护的紧迫性,并推动负责任AI落地的讨论。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) represent the cutting edge of AI technology, with DeepSeek models emerging as a leading open-source alternative offering competitive performance to closed-source systems. While these models demonstrate remarkable capabilities, their vision-language integration mechanisms introduce specific vulnerabilities. We implement an adapted embedding manipulation attack on DeepSeek Janus that induces targeted visual hallucinations through systematic optimization of image embeddings. Through extensive experimentation across COCO, DALL-E 3, and SVIT datasets, we achieve hallucination rates of up to 98.0% while maintaining high visual fidelity (SSIM > 0.88) of the manipulated images on open-ended questions. Our analysis demonstrates that both 1B and 7B variants of DeepSeek Janus are susceptible to these attacks, with closed-form evaluation showing consistently higher hallucination rates compared to open-ended questioning. We introduce a novel multi-prompt hallucination detection framework using LLaMA-3.1 8B Instruct for robust evaluation. The implications of these findings are particularly concerning given DeepSeek's open-source nature and widespread deployment potential. This research emphasizes the critical need for embedding-level security measures in MLLM deployment pipelines and contributes to the broader discussion of responsible AI implementation.

多模态模型安全幻觉生成嵌入攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。