arXiv:2502.18518cs.CRcs.AI2025-02被引 2

研究发现大模型对局部数据投毒攻击极脆弱,且压缩后更易受攻击。

Swallowing the Poison Pills: Insights from Vulnerability Disparity Among LLMs

  • 通过修改时间、空间等信息诱导模型局部记忆失效,几乎不影响常规性能。
  • 长尾知识错误率提升54.6%,压缩模型错误率最高升25.5%。
  • 揭示模型压缩中的安全-效率矛盾,适合关注模型安全的开发者参考。

现代大语言模型易受‘毒丸攻击’影响:局部数据污染可改变特定事实知识,同时保持整体模型性能。我们系统性验证此类攻击利用了模型固有架构特性,在长尾知识上检索错误率提升54.6%,压缩模型相比原模型错误率最高上升25.5%。通过控制变量(如时间、空间、实体修改),该方法引发局部记忆退化,对标准基准测试影响极小(如MMLU/GPQA性能下降<2%),具备隐蔽性。研究发现:(1) 长尾知识脆弱性源于参数冗余不足;(2) 模型压缩扩大攻击面,剪枝/蒸馏模型仅需30%更少毒样本即可造成同等损害;(3) 关联记忆导致损伤扩散与叠加,尤其影响主流话题。这些结果警示当前扩展范式——攻击成本降低而防御复杂度上升。本工作将毒丸攻击视为安全威胁与诊断工具,揭示语言模型压缩中的关键安全-效率权衡,挑战现有安全假设。

原文摘要 · Abstract (English)

Modern large language models (LLMs) exhibit critical vulnerabilities to poison pill attacks: localized data poisoning that alters specific factual knowledge while preserving overall model utility. We systematically demonstrate these attacks exploit inherent architectural properties of LLMs, achieving 54.6% increased retrieval inaccuracy on long-tail knowledge versus dominant topics and up to 25.5% increase retrieval inaccuracy on compressed models versus original architectures. Through controlled mutations (e.g., temporal/spatial/entity alterations) and, our method induces localized memorization deterioration with negligible impact on models' performance on regular standard benchmarks (e.g., <2% performance drop on MMLU/GPQA), leading to potential detection evasion. Our findings suggest: (1) Disproportionate vulnerability in long-tail knowledge may result from reduced parameter redundancy; (2) Model compression may increase attack surfaces, with pruned/distilled models requiring 30% fewer poison samples for equivalent damage; (3) Associative memory enables both spread of collateral damage to related concepts and amplification of damage from simultaneous attack, particularly for dominant topics. These findings raise concerns over current scaling paradigms since attack costs are lowering while defense complexity is rising. Our work establishes poison pills as both a security threat and diagnostic tool, revealing critical security-efficiency trade-offs in language model compression that challenges prevailing safety assumptions.

模型安全数据投毒压缩风险长尾知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。