首个针对蛋白质大模型的红队测试框架,发现其存在生物安全风险
SafeProtein: Red-Teaming Framework and Benchmark for Protein Foundation Models
- 通过多模态提示工程与启发式搜索生成对抗性蛋白质序列
- 在ESM3模型上实现最高70%的攻击成功率,暴露安全漏洞
- 适合关注生物安全与大模型防护的研究者和开发者
蛋白质几乎参与所有生物过程。深度学习的进步推动了蛋白质基础模型的发展,在理解与设计蛋白质方面取得显著进展。然而,这些模型缺乏系统的红队测试,可能被滥用生成具有生物安全风险的蛋白质。本文提出SafeProtein,据我们所知是首个专为蛋白质基础模型设计的红队框架。该框架结合多模态提示工程与启发式束搜索,系统化地设计红队方法并测试蛋白质基础模型。同时构建了SafeProtein-Bench,包含人工标注的红队基准数据集和全面评估协议。实验显示,SafeProtein可在最先进的蛋白质基础模型(如ESM3)上持续实现越狱攻击,最高成功率达70%,揭示当前模型潜在的生物安全风险,并为前沿模型的安全防护技术提供重要参考。代码将公开于https://github.com/jigang-fan/SafeProtein。
原文摘要 · Abstract (English)
Proteins play crucial roles in almost all biological processes. The advancement of deep learning has greatly accelerated the development of protein foundation models, leading to significant successes in protein understanding and design. However, the lack of systematic red-teaming for these models has raised serious concerns about their potential misuse, such as generating proteins with biological safety risks. This paper introduces SafeProtein, the first red-teaming framework designed for protein foundation models to the best of our knowledge. SafeProtein combines multimodal prompt engineering and heuristic beam search to systematically design red-teaming methods and conduct tests on protein foundation models. We also curated SafeProtein-Bench, which includes a manually constructed red-teaming benchmark dataset and a comprehensive evaluation protocol. SafeProtein achieved continuous jailbreaks on state-of-the-art protein foundation models (up to 70% attack success rate for ESM3), revealing potential biological safety risks in current protein foundation models and providing insights for the development of robust security protection technologies for frontier models. The codes will be made publicly available at https://github.com/jigang-fan/SafeProtein.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。