系统梳理大模型提示词攻击三类手法,提升安全评估精度
SoK: Prompt Hacking of Large Language Models
- 区分越狱、泄露、注入三类提示攻击,厘清其异同
- 提出五分类响应评估框架,替代传统二元判断
- 适合关注大模型安全与防御的研究者和开发者
基于大语言模型的应用安全与鲁棒性仍是人工智能领域的重要挑战。其中,提示词攻击是关键威胁之一,可能严重破坏大模型系统的安全性和可靠性。本文对三类典型的提示词攻击——越狱、泄露与注入进行了全面而系统的梳理,尽管它们存在重叠特征,但本质差异显著。为提升大模型应用的评估能力,我们提出一种新框架,将大模型输出分为五类,突破传统二元分类局限,提供更精细的行为洞察,有助于精准诊断问题并针对性增强系统安全性与鲁棒性。
原文摘要 · Abstract (English)
The safety and robustness of large language models (LLMs) based applications remain critical challenges in artificial intelligence. Among the key threats to these applications are prompt hacking attacks, which can significantly undermine the security and reliability of LLM-based systems. In this work, we offer a comprehensive and systematic overview of three distinct types of prompt hacking: jailbreaking, leaking, and injection, addressing the nuances that differentiate them despite their overlapping characteristics. To enhance the evaluation of LLM-based applications, we propose a novel framework that categorizes LLM responses into five distinct classes, moving beyond the traditional binary classification. This approach provides more granular insights into the AI's behavior, improving diagnostic precision and enabling more targeted enhancements to the system's safety and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。