构建对抗性提示框架,评估生成式AI安全漏洞。
Adversarial Prompting Framework for AI Safety Assessment

- 分层级生成对抗提示,从直接攻击到编码隐匿攻击
- 编码类攻击绕过安全机制成功率最高,暴露显著漏洞
- 适合企业部署的自动化安全评估,提供量化指标
近年来,人工智能尤其是生成式AI在各行业应用迅速增长。然而,这些模型也可能因恶意攻击者发起的对抗性提示攻击(APA)而面临新型网络安全威胁。本文实现了一种对抗性提示框架(APF),用于全面评估AI安全性。该框架通过生成多层次、结构化的对抗性提示,系统测试模型在不同攻击复杂度下的鲁棒性,涵盖从直接有害请求到高级编码攻击。实验表明,在多种攻击向量中,编码类提示的成功率最高,显著暴露了模型的安全缺陷。该方法已在企业环境中验证,具备自动化测试能力与量化评估指标,为实际部署提供安全参考。
原文摘要 · Abstract (English)
Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use of these models may also expose systems to new forms of cyberattacks by different malicious actors -- adversarial prompt attack (APA) being one of the most prominent examples of such threats. This paper presents the implementation of an Adversarial Prompting Framework (APF) for a comprehensive assessment of AI safety. The framework systematically evaluates the resilience of the AI model through the generation of structured adversarial prompts at multiple sophistication levels, from direct harmful requests to advanced encoding-based attacks. Our implementation demonstrates the practical application of this methodology in enterprise environments, providing automated testing capabilities with quantitative security assessment metrics. The results indicate significant variations in the model vulnerabilities across different attack vectors, with encoded prompts presenting the highest success rates in bypassing safety mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。