开发开源工具PyRIT,帮助发现生成式AI的潜在风险与越狱漏洞。
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
- 构建可组合的通用框架,支持多模态模型的安全测试
- 能识别新型危害与越狱攻击,提升红队测试效率
- 适合安全研究人员和AI开发者用于风险评估
生成式人工智能(GenAI)正广泛融入日常生活。随着算力和数据的增长,单模态与多模态模型大量涌现。随着生态成熟,对可扩展、模型无关的风险识别框架的需求日益迫切。为此,我们提出Python风险识别工具包(PyRIT),一个开源框架,旨在增强生成式AI系统的红队测试能力。PyRIT是模型与平台无关的工具,可探测并识别多模态生成式AI模型中的新型危害、风险及越狱行为。其可组合架构支持核心组件复用,并可扩展至未来模型与模态。本文详述了生成式AI红队测试的独特挑战,PyRIT的设计与功能,以及在真实场景中的应用实例。
原文摘要 · Abstract (English)
Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives. The increase in computational power and data availability has led to a proliferation of both single- and multi-modal models. As the GenAI ecosystem matures, the need for extensible and model-agnostic risk identification frameworks is growing. To meet this need, we introduce the Python Risk Identification Toolkit (PyRIT), an open-source framework designed to enhance red teaming efforts in GenAI systems. PyRIT is a model- and platform-agnostic tool that enables red teamers to probe for and identify novel harms, risks, and jailbreaks in multimodal generative AI models. Its composable architecture facilitates the reuse of core building blocks and allows for extensibility to future models and modalities. This paper details the challenges specific to red teaming generative AI systems, the development and features of PyRIT, and its practical applications in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。