arXiv:2509.05883cs.CRcs.AI2025-09被引 11

测试8款商用大模型,发现多模态提示注入漏洞普遍存在。

Multimodal Prompt Injection Attacks: Risks and Defenses for Modern LLMs

  • 设计四类攻击:直接、间接、图像型和提示泄露
  • 所有模型均被攻破,仅Claude 3表现稍强
  • 建议加入输入归一化等额外防护措施

大型语言模型(LLMs)近年来迅速普及,被广泛应用于咨询、信息检索等多个领域。尽管其能精准理解指令并生成类人回复,但部署广泛也带来了显著安全风险,尤其是提示注入与越狱攻击。为系统评估大模型对外部提示注入的脆弱性,我们对八款商用模型进行了实验,均未启用额外清洗,仅依赖内置防护。结果揭示出可被利用的缺陷,强调亟需更强的安全机制。研究考察了四类攻击:直接注入、间接(外部)注入、基于图像的注入以及提示泄露。对比分析显示,Claude 3表现出相对更高的鲁棒性;然而实证结果证实,仍需通过输入归一化等附加防御手段实现可靠保护。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have seen rapid adoption in recent years, with industries increasingly relying on them to maintain a competitive advantage. These models excel at interpreting user instructions and generating human-like responses, leading to their integration across diverse domains, including consulting and information retrieval. However, their widespread deployment also introduces substantial security risks, most notably in the form of prompt injection and jailbreak attacks. To systematically evaluate LLM vulnerabilities -- particularly to external prompt injection -- we conducted a series of experiments on eight commercial models. Each model was tested without supplementary sanitization, relying solely on its built-in safeguards. The results exposed exploitable weaknesses and emphasized the need for stronger security measures. Four categories of attacks were examined: direct injection, indirect (external) injection, image-based injection, and prompt leakage. Comparative analysis indicated that Claude 3 demonstrated relatively greater robustness; nevertheless, empirical findings confirm that additional defenses, such as input normalization, remain necessary to achieve reliable protection.

提示注入大模型安全多模态攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。