InjectLab构建了针对大模型的对抗性攻击框架,系统梳理25种提示层威胁。
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
- 基于MITRE ATT&CK思想,构建六战术25种提示攻击技术矩阵
- 涵盖指令覆盖、身份冒充等真实攻击手法,附检测与防御策略
- 开源可执行测试工具,适合安全研究者与模型开发者使用
大型语言模型正在重塑人机交互方式,但其也面临提示层攻击的新风险。本文提出InjectLab,一个结构化、开源的对抗性威胁建模框架,借鉴MITRE ATT&CK理念,聚焦提示层的恶意行为。该框架包含超过25种技术,按六类核心战术组织,涵盖指令覆盖、身份替换、多智能体滥用等攻击手段。每项技术均配有检测建议、缓解方案及基于YAML的仿真测试用例,配套Python工具支持快速执行。论文阐述了框架设计,对比现有AI威胁分类体系,并展望其作为社区驱动的安全基石的未来发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are changing the way people interact with technology. Tools like ChatGPT and Claude AI are now common in business, research, and everyday life. But with that growth comes new risks, especially prompt-based attacks that exploit how these models process language. InjectLab is a security framework designed to address that problem. This paper introduces InjectLab as a structured, open-source matrix that maps real-world techniques used to manipulate LLMs. The framework is inspired by MITRE ATT&CK and focuses specifically on adversarial behavior at the prompt layer. It includes over 25 techniques organized under six core tactics, covering threats like instruction override, identity swapping, and multi-agent exploitation. Each technique in InjectLab includes detection guidance, mitigation strategies, and YAML-based simulation tests. A Python tool supports easy execution of prompt-based test cases. This paper outlines the framework's structure, compares it to other AI threat taxonomies, and discusses its future direction as a practical, community-driven foundation for securing language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。