arXiv:2506.06518cs.CRcs.LG2025-06综述被引 14

系统梳理大模型中毒攻击,构建通用分类框架。

A Systematic Review of Poisoning Attacks Against Large Language Models

  • 提出四维度攻击分类与六项评估指标
  • 归纳四类典型中毒攻击:概念、隐蔽、持久、任务专属
  • 统一术语体系,助研究者厘清安全风险

随着预训练大语言模型(LLMs)及其训练数据的广泛可用,其使用带来的安全风险日益受到关注。其中一类风险是大模型中毒攻击,即攻击者通过篡改部分训练过程,使模型产生恶意行为。作为新兴研究方向,现有框架和术语多源自早期分类中毒文献,难以适配生成式大模型场景。本文对已发表的LLM中毒攻击进行系统性综述,澄清安全影响,并解决文献中术语不一致的问题。提出一个适用于多种攻击类型的综合中毒威胁模型,包含四个攻击规格(定义攻击策略与实施方式)和六个评估指标(衡量攻击关键特征)。基于该框架,从四个关键维度组织讨论:概念中毒、隐蔽中毒、持久中毒和特定任务中毒,以更清晰地理解当前安全风险全景。

原文摘要 · Abstract (English)

With the widespread availability of pretrained Large Language Models (LLMs) and their training datasets, concerns about the security risks associated with their usage has increased significantly. One of these security risks is the threat of LLM poisoning attacks where an attacker modifies some part of the LLM training process to cause the LLM to behave in a malicious way. As an emerging area of research, the current frameworks and terminology for LLM poisoning attacks are derived from earlier classification poisoning literature and are not fully equipped for generative LLM settings. We conduct a systematic review of published LLM poisoning attacks to clarify the security implications and address inconsistencies in terminology across the literature. We propose a comprehensive poisoning threat model applicable to categorize a wide range of LLM poisoning attacks. The poisoning threat model includes four poisoning attack specifications that define the logistics and manipulation strategies of an attack as well as six poisoning metrics used to measure key characteristics of an attack. Under our proposed framework, we organize our discussion of published LLM poisoning literature along four critical dimensions of LLM poisoning attacks: concept poisons, stealthy poisons, persistent poisons, and poisons for unique tasks, to better understand the current landscape of security risks.

大模型安全中毒攻击威胁建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。