黑客通过恶意字体隐藏指令,让大模型偷偷执行攻击
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
- 用恶意字体篡改字符映射,藏起攻击代码
- 利用外部资源绕过安全机制,成功率达70%以上
- 适合关注大模型安全的开发者与研究人员
大型语言模型(LLMs)正越来越多地集成实时网络搜索功能及模型上下文协议(MCP)。这种扩展可能引入新的安全漏洞。我们系统性研究了通过网页等外部资源中的恶意字体注入隐蔽对抗性提示对LLM的威胁。攻击者通过操纵代码到字形的映射关系,注入用户不可见的欺骗性内容。评估了两种关键攻击场景:(1) 恶意内容中继,(2) 通过MCP工具引发敏感信息泄露。实验表明,通过外部资源注入的间接提示可绕过LLM安全机制,成功率因数据敏感性和提示设计而异,最高可达70%以上。本研究凸显了在处理外部内容时加强大模型部署安全措施的紧迫性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly equipped with capabilities of real-time web search and integrated with protocols like Model Context Protocol (MCP). This extension could introduce new security vulnerabilities. We present a systematic investigation of LLM vulnerabilities to hidden adversarial prompts through malicious font injection in external resources like webpages, where attackers manipulate code-to-glyph mapping to inject deceptive content which are invisible to users. We evaluate two critical attack scenarios: (1) "malicious content relay" and (2) "sensitive data leakage" through MCP-enabled tools. Our experiments reveal that indirect prompts with injected malicious font can bypass LLM safety mechanisms through external resources, achieving varying success rates based on data sensitivity and prompt design. Our research underscores the urgent need for enhanced security measures in LLM deployments when processing external content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。