提出一种高效文本注入攻击,可低成本误导大视觉语言模型。
Text Prompt Injection of Vision Language Models
- 设计轻量级文本提示注入算法,无需高算力
- 在大模型上实现高成功率误导,效果优于现有方法
- 适合研究模型安全与对抗攻击的学者参考
大型视觉语言模型的广泛应用引发了显著的安全隐患。本项目研究了一种简单而有效的文本提示注入攻击方法,旨在误导这些模型。我们开发了一种针对此类攻击的算法,并通过实验验证了其有效性和高效性。相较于其他攻击方法,该方法对大型模型尤为有效,且对计算资源需求极低,可在资源受限环境下实现对大模型的可靠操控。
原文摘要 · Abstract (English)
The widespread application of large vision language models has significantly raised safety concerns. In this project, we investigate text prompt injection, a simple yet effective method to mislead these models. We developed an algorithm for this type of attack and demonstrated its effectiveness and efficiency through experiments. Compared to other attack methods, our approach is particularly effective for large models without high demand for computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。