提出可自动检测大模型越狱漏洞的开源安全评估框架
AVISE: Framework for Evaluating the Security of AI Systems

- 构建模块化框架AVISE,支持自动化安全测试
- 测试集准确率达92%,能发现各类大模型的越狱漏洞
- 适合研究人员和从业者进行可复现的安全评测
随着人工智能系统在关键领域广泛应用,其安全漏洞带来的高风险攻击和系统故障日益突出。然而,系统化的AI安全评估方法仍不成熟。本文提出AVISE(AI漏洞识别与安全评估)框架,一个模块化开源工具,用于识别和评估AI系统及模型的安全性。作为演示,我们将基于心智理论的多轮红皇后攻击扩展为对抗语言模型(ALM)增强攻击,并开发了自动化安全评估测试(SET),用于发现语言模型中的越狱漏洞。SET包含25个测试用例和一个评估语言模型(ELM),可判断每个用例是否成功越狱目标模型,达到92%准确率、0.91的F1分数和0.83的马修相关系数。我们使用SET评估了九个不同规模的近期发布语言模型,发现它们均不同程度地易受增强型红皇后攻击。AVISE为研究者和产业界提供了可扩展的自动化测试基础,推动更严格、可复现的AI安全评估。
原文摘要 · Abstract (English)
As artificial intelligence (AI) systems are increasingly deployed across critical domains, their security vulnerabilities pose growing risks of high-profile exploits and consequential system failures. Yet systematic approaches to evaluating AI security remain underdeveloped. In this paper, we introduce AVISE (AI Vulnerability Identification and Security Evaluation), a modular open-source framework for identifying vulnerabilities in and evaluating the security of AI systems and models. As a demonstration of the framework, we extend the theory-of-mind-based multi-turn Red Queen attack into an Adversarial Language Model (ALM) augmented attack and develop an automated Security Evaluation Test (SET) for discovering jailbreak vulnerabilities in language models. The SET comprises 25 test cases and an Evaluation Language Model (ELM) that determines whether each test case was able to jailbreak the target model, achieving 92% accuracy, an F1-score of 0.91, and a Matthews correlation coefficient of 0.83. We evaluate nine recently released language models of diverse sizes with the SET and find that all are vulnerable to the augmented Red Queen attack to varying degrees. AVISE provides researchers and industry practitioners with an extensible foundation for developing and deploying automated SETs, offering a concrete step toward more rigorous and reproducible AI security evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。