arXiv:2505.09974cs.CRcs.AI2025-05被引 2

用伪造的网络安全数据微调大模型会显著降低其安全性。

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

  • 用伪恶意数据微调模型,导致安全防护能力下降。
  • 微调后模型在提示注入攻击下失败率从9.1%升至68.7%。
  • 通过重写指令对齐安全策略,可提升安全性且不影响性能。

大型语言模型(LLMs)已广泛应用于网络安全领域,虽能增强威胁分析与恶意软件检测能力,但也带来个人数据泄露和自动生成新恶意软件等安全风险。本文基于最新研究,发现使用伪恶意网络安全数据微调模型会严重削弱其安全性,为此采用 garak 红队测试框架与 OWASP Top 10 for LLM Applications,评估了 Mistral 7B、Llama 3 8B、Gemma 2 9B 与 DeepSeek R1 8B 四个开源模型。结果验证并扩展了先前结论:所有模型的安全韧性均下降,例如 Mistral 7B 在提示注入攻击下的失败率由 9.1% 上升至 68.7%。本文提出一种新型安全对齐方法,通过重构指令-响应对以加入明确安全提示与伦理考量,实证表明该方法可在保持技术效用的同时有效提升安全性,为构建更安全、可信、伦理对齐的 LLM 提供可行路径。

原文摘要 · Abstract (English)

Large language models (LLMs) have been used in many application domains, including cyber security. The application of LLMs in the cyber security domain presents significant opportunities, such as for enhancing threat analysis and malware detection, but it can also introduce critical risks and safety concerns, including potential personal data leakage and automated generation of new malware. Building on recent findings that fine-tuning LLMs with pseudo-malicious cyber security data significantly compromises their safety, this paper presents a comprehensive validation and extension of these safety risks using a different evaluation framework. We employ the garak red teaming framework with the OWASP Top 10 for LLM Applications to assess four open-source LLMs: Mistral 7B, Llama 3 8B, Gemma 2 9B, and DeepSeek R1 8B. Our evaluation confirms and extends previous findings, showing that fine-tuning reduces safety resilience across all tested LLMs (e.g., the failure rate of Mistral 7B against prompt injection increases from 9.1% to 68.7%). We further propose and evaluate a novel safety alignment approach that carefully rewords instruction-response pairs to include explicit safety precautions and ethical considerations. This work validates previous safety concerns through independent evaluation and introduces new methods for mitigating these risks, contributing towards the development of secure, trustworthy, and ethically aligned LLMs. This approach demonstrates that it is possible to maintain or even improve model safety while preserving technical utility, offering a practical path towards developing safer fine-tuning methodologies.

大模型安全微调风险红队测试伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。