arXiv:2503.09334cs.CRcs.AI2025-03被引 8

构建安全风险数据集,揭示网络安全LLM微调中的性能与安全权衡。

CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning

  • 构建5.49万条伪恶意指令数据,覆盖漏洞分析等安全任务。
  • 微调后模型准确率最高达92.5%,但抗注入攻击能力下降至0.15。
  • 适合关注安全敏感场景下LLM鲁棒性的研究人员和开发者。

大型语言模型(LLMs)在网络安全应用中既带来机遇也引发安全风险。本文提出CyberLLMInstruct,一个包含54,928条伪恶意指令-响应对的数据集,涵盖恶意软件分析、钓鱼模拟和零日漏洞等任务。基于七个开源LLM的全面评估显示:微调虽提升任务性能(在CyberMetric上最高达92.50%准确率),却严重削弱所有模型在各类攻击向量下的安全韧性(如Llama 3.1 8B在提示注入攻击下的安全得分从0.95降至0.15)。数据集融合CTF挑战、学术论文、行业报告及CVE数据库,确保领域覆盖全面。研究揭示了在对抗性环境中保障LLM安全的独特挑战,强调需发展兼顾性能与安全的微调方法。

原文摘要 · Abstract (English)

The integration of large language models (LLMs) into cyber security applications presents both opportunities and critical safety risks. We introduce CyberLLMInstruct, a dataset of 54,928 pseudo-malicious instruction-response pairs spanning cyber security tasks including malware analysis, phishing simulations, and zero-day vulnerabilities. Our comprehensive evaluation using seven open-source LLMs reveals a critical trade-off: while fine-tuning improves cyber security task performance (achieving up to 92.50% accuracy on CyberMetric), it severely compromises safety resilience across all tested models and attack vectors (e.g., Llama 3.1 8B's security score against prompt injection drops from 0.95 to 0.15). The dataset incorporates diverse sources including CTF challenges, academic papers, industry reports, and CVE databases to ensure comprehensive coverage of cyber security domains. Our findings highlight the unique challenges of securing LLMs in adversarial domains and establish the critical need for developing fine-tuning methodologies that balance performance gains with safety preservation in security-sensitive domains.

网络安全LLM安全数据集微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。