arXiv:2410.15236cs.CRcs.AI2024-10被引 45

系统梳理大模型漏洞与攻防策略,助力安全落地

Jailbreaking and Mitigation of Vulnerabilities in Large Language Models

  • 按提示、模型、多模态等分类攻击手段,覆盖越狱与后门
  • 总结过滤、对齐、多代理等防御方法优劣与适用场景
  • 适合关注AI安全、模型防护的研究者与开发者

大型语言模型(LLMs)在自然语言理解与生成方面取得显著进展,推动医疗、软件工程及对话系统等领域的应用。然而近年来,其在提示注入和越狱攻击下暴露出明显漏洞。本文综述该领域研究现状,将攻击方式大致分为基于提示、基于模型、多模态及多语言四类,涵盖对抗性提示、后门注入与跨模态攻击等技术。同时梳理了提示过滤、转换、对齐、多智能体防御与自调节等防御机制,评估其优势与局限。讨论了评估模型安全与鲁棒性的关键指标与基准测试,指出交互环境下攻击成功率量化困难及数据集偏差等问题。识别当前研究空白,提出未来方向:强化对齐策略、应对演化攻击的先进防御、越狱检测自动化以及伦理与社会影响考量。强调需持续合作以提升模型安全性,确保安全部署。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have transformed artificial intelligence by advancing natural language understanding and generation, enabling applications across fields beyond healthcare, software engineering, and conversational systems. Despite these advancements in the past few years, LLMs have shown considerable vulnerabilities, particularly to prompt injection and jailbreaking attacks. This review analyzes the state of research on these vulnerabilities and presents available defense strategies. We roughly categorize attack approaches into prompt-based, model-based, multimodal, and multilingual, covering techniques such as adversarial prompting, backdoor injections, and cross-modality exploits. We also review various defense mechanisms, including prompt filtering, transformation, alignment techniques, multi-agent defenses, and self-regulation, evaluating their strengths and shortcomings. We also discuss key metrics and benchmarks used to assess LLM safety and robustness, noting challenges like the quantification of attack success in interactive contexts and biases in existing datasets. Identifying current research gaps, we suggest future directions for resilient alignment strategies, advanced defenses against evolving attacks, automation of jailbreak detection, and consideration of ethical and societal impacts. This review emphasizes the need for continued research and cooperation within the AI community to enhance LLM security and ensure their safe deployment.

大模型安全越狱攻击防御机制综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。