arXiv:2511.02600cs.CRcs.AI2025-11

恶意数据污染可让安全AI误判真实告警,导致防护失效。

On The Dangers of Poisoned LLMs In Security Automation

  • 用少量恶意数据微调大模型,植入定向偏见。
  • 攻击使模型持续忽略特定用户的真实安全告警。
  • 适合关注AI安全风险与可信部署的研究者。

本文研究了'大模型污染'带来的风险,即在训练过程中有意或无意引入恶意或偏见数据。我们证明,一个看似性能提升的微调后大模型,可能引入显著偏差,导致基于LLM的告警调查系统完全被绕过——当提示利用该偏差时,系统会忽略真实告警。使用微调后的Llama3.1 8B和Qwen3 4B模型,我们展示了针对性污染攻击如何使模型持续忽视来自特定用户的真实正例告警。此外,论文提出了若干缓解策略与最佳实践,以提升安全应用中大模型的可信度、鲁棒性并降低风险。

原文摘要 · Abstract (English)

This paper investigates some of the risks introduced by "LLM poisoning," the intentional or unintentional introduction of malicious or biased data during model training. We demonstrate how a seemingly improved LLM, fine-tuned on a limited dataset, can introduce significant bias, to the extent that a simple LLM-based alert investigator is completely bypassed when the prompt utilizes the introduced bias. Using fine-tuned Llama3.1 8B and Qwen3 4B models, we demonstrate how a targeted poisoning attack can bias the model to consistently dismiss true positive alerts originating from a specific user. Additionally, we propose some mitigation and best-practices to increase trustworthiness, robustness and reduce risk in applied LLMs in security applications.

AI安全模型污染安全自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。