压缩投影器严重威胁大视觉语言模型安全,未压缩的则更可靠。
The Security Threat of Compressed Projectors in Large Vision-Language Models
- 对比压缩与未压缩投影器的安全性差异
- 压缩投影器在极少信息下即可被攻破
- 适合关注模型安全性的研究人员参考
选择合适的视觉语言投影器(VLP)对大视觉语言模型(LVLM)的成功训练至关重要。主流VLP可分为压缩与未压缩两类,各自在性能与计算效率上各有优势。然而其安全影响尚未充分研究。我们全面评估发现二者安全特性显著不同:压缩投影器存在严重漏洞,攻击者仅需少量结构信息即可成功破坏LVLM;而未压缩投影器表现出强安全性,未引入额外风险。该结果为研究人员选择更安全的VLP提供了关键指导。代码已公开于https://github.com/btzyd/TCP。
原文摘要 · Abstract (English)
The choice of a suitable visual language projector (VLP) is critical to the successful training of large visual language models (LVLMs). Mainstream VLPs can be broadly categorized into compressed and uncompressed projectors, and each offers distinct advantages in performance and computational efficiency. However, their security implications have not been thoroughly examined. Our comprehensive evaluation reveals significant differences in their security profiles: compressed projectors exhibit substantial vulnerabilities, allowing adversaries to successfully compromise LVLMs even with minimal knowledge of structure information. In stark contrast, uncompressed projectors demonstrate robust security properties and do not introduce additional vulnerabilities. These findings provide critical guidance for researchers in selecting optimal VLPs that enhance the security and reliability of visual language models. The code is available at https://github.com/btzyd/TCP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。