用化学分子式符号绕过模型安全限制,诱导生成危险合成方案。
SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis
- 用SMILES分子标识符作为提示,隐蔽触发模型生成有害指令。
- 该方法在测试中成功绕过主流安全机制,攻击成功率显著高于传统方式。
- 适用于研究模型安全漏洞的学者,或需防范化学滥用的AI开发者。
大型语言模型(LLMs)在各领域的广泛应用引发了对其传播危险信息的担忧。本文聚焦化学领域中LLM的安全漏洞,探讨其生成危险物质合成路径的能力。我们评估了红队测试、显式提示和隐式提示等多种注入攻击方法,并提出一种新型攻击技术——SMILES-prompting,利用简化分子输入线性系统(SMILES)来指代化学物质。实验表明,该方法能有效规避现有安全机制。研究凸显了亟需加强领域特定防护措施,以防止模型被滥用,提升其社会正向价值。
原文摘要 · Abstract (English)
The increasing integration of large language models (LLMs) across various fields has heightened concerns about their potential to propagate dangerous information. This paper specifically explores the security vulnerabilities of LLMs within the field of chemistry, particularly their capacity to provide instructions for synthesizing hazardous substances. We evaluate the effectiveness of several prompt injection attack methods, including red-teaming, explicit prompting, and implicit prompting. Additionally, we introduce a novel attack technique named SMILES-prompting, which uses the Simplified Molecular-Input Line-Entry System (SMILES) to reference chemical substances. Our findings reveal that SMILES-prompting can effectively bypass current safety mechanisms. These findings highlight the urgent need for enhanced domain-specific safeguards in LLMs to prevent misuse and improve their potential for positive social impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。