测试大模型在偏好学习中对数据投毒的脆弱性,发现模型越大越不安全。
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
- 设计两类攻击在8个真实场景下测试21个模型
- 模型参数量越大,抗投毒能力反而越弱
- 投毒效果可泛化到未见过的触发词,威胁隐蔽
偏好学习是当前大语言模型对齐的核心环节,但这一过程易受数据投毒攻击影响。为评估此类风险,我们提出PoisonBench,一个用于衡量大语言模型在偏好学习阶段对数据投毒敏感性的基准。数据投毒可使模型输出隐藏恶意内容或偏见,在表面正常运作的同时生成有害或非预期结果。我们在八个真实场景中部署两种攻击类型,评估了21个广泛使用的模型。结果显示:(1) 参数量增大并未提升对投毒攻击的抵抗力;(2) 攻击效果与投毒比例呈对数线性关系;(3) 投毒影响可泛化至未包含在污染数据中的外推触发词。这些发现揭示了现有偏好学习技术的弱点,凸显了亟需更鲁棒的防御机制以应对恶意模型和数据操控。
原文摘要 · Abstract (English)
Preference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBench, a benchmark for evaluating large language models' susceptibility to data poisoning during preference learning. Data poisoning attacks can manipulate large language model responses to include hidden malicious content or biases, potentially causing the model to generate harmful or unintended outputs while appearing to function normally. We deploy two distinct attack types across eight realistic scenarios, assessing 21 widely-used models. Our findings reveal concerning trends: (1) Scaling up parameter size does not inherently enhance resilience against poisoning attacks; (2) There exists a log-linear relationship between the effects of the attack and the data poison ratio; (3) The effect of data poisoning can generalize to extrapolated triggers that are not included in the poisoned data. These results expose weaknesses in current preference learning techniques, highlighting the urgent need for more robust defenses against malicious models and data manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。