LLM在网络安全情报任务中表现不可靠,易误判且盲目自信。
Large Language Models Are Unreliable for Cyber Threat Intelligence
- 测试三种顶级LLM在零样本、少样本和微调下的表现
- 在350份报告上发现性能不足、结果不一致、自信过度
- 适合关注AI安全风险的从业者,不建议直接用于真实威胁研判
近期多项研究主张大型语言模型(LLMs)可缓解网络安全领域数据过载问题,提升网络威胁情报(CTI)任务自动化水平。本文提出一种评估方法,不仅支持在零样本、少样本和微调条件下测试LLM在CTI任务中的表现,还可量化其一致性与置信度。我们对三种前沿LLM在包含350份威胁情报报告的数据集上进行了实验,揭示了依赖LLM进行CTI可能带来的潜在安全风险。结果表明,这些模型在处理真实规模报告时无法保证足够性能,同时表现出显著的不一致性和过度自信。少样本学习与微调仅部分改善结果,这使得在缺乏标注数据且置信度至关重要的实际CTI场景中使用LLM的可行性受到质疑。
原文摘要 · Abstract (English)
Several recent works have argued that Large Language Models (LLMs) can be used to tame the data deluge in the cybersecurity field, by improving the automation of Cyber Threat Intelligence (CTI) tasks. This work presents an evaluation methodology that other than allowing to test LLMs on CTI tasks when using zero-shot learning, few-shot learning and fine-tuning, also allows to quantify their consistency and their confidence level. We run experiments with three state-of-the-art LLMs and a dataset of 350 threat intelligence reports and present new evidence of potential security risks in relying on LLMs for CTI. We show how LLMs cannot guarantee sufficient performance on real-size reports while also being inconsistent and overconfident. Few-shot learning and fine-tuning only partially improve the results, thus posing doubts about the possibility of using LLMs for CTI scenarios, where labelled datasets are lacking and where confidence is a fundamental factor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。