发现大模型在自动算法设计中存在高危安全漏洞,可被恶意指令攻破。
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
- 构建恶意优化算法请求基准集MalOptBench,测试模型对诱导性指令的响应。
- 13个主流模型平均攻击成功率83.59%,原始有害提示危害评分达4.28/5。
- 现有防护机制效果有限,尤其对新型定制化越狱方法防御力不足。
大语言模型(LLMs)的广泛应用引发了对其滥用风险与安全问题的日益关注。尽管已有研究探讨了通用场景、代码生成及基于代理的应用中的安全性,但自动化算法设计方面的漏洞仍被忽视。本文聚焦智能优化算法设计这一关键领域,提出MalOptBench基准,包含60个恶意优化算法请求,并设计针对性越狱方法MOBjailbreak。对包括最新GPT-5和DeepSeek-V3.1在内的13个主流模型进行评估,结果显示大多数模型仍高度易受攻击:原始有害提示下平均攻击成功率达83.59%,平均危害评分为4.28/5;在MOBjailbreak下几乎完全失效。同时评估了现成的即插即用式防御方案,发现其对MOBjailbreak仅具微弱防护能力,且易引发过度保守行为。研究凸显了强化对齐技术以防范算法设计场景下模型滥用的紧迫性。
原文摘要 · Abstract (English)
The widespread deployment of large language models (LLMs) has raised growing concerns about their misuse risks and associated safety issues. While prior studies have examined the safety of LLMs in general usage, code generation, and agent-based applications, their vulnerabilities in automated algorithm design remain underexplored. To fill this gap, this study investigates this overlooked safety vulnerability, with a particular focus on intelligent optimization algorithm design, given its prevalent use in complex decision-making scenarios. We introduce MalOptBench, a benchmark consisting of 60 malicious optimization algorithm requests, and propose MOBjailbreak, a jailbreak method tailored for this scenario. Through extensive evaluation of 13 mainstream LLMs including the latest GPT-5 and DeepSeek-V3.1, we reveal that most models remain highly susceptible to such attacks, with an average attack success rate of 83.59% and an average harmfulness score of 4.28 out of 5 on original harmful prompts, and near-complete failure under MOBjailbreak. Furthermore, we assess state-of-the-art plug-and-play defenses that can be applied to closed-source models, and find that they are only marginally effective against MOBjailbreak and prone to exaggerated safety behaviors. These findings highlight the urgent need for stronger alignment techniques to safeguard LLMs against misuse in algorithm design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。