用大模型提升威胁优先级,发现三大真实漏洞。
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
- 分层+自回归+人工监督,系统分析大模型在安全中的失败原因。
- 实测发现三类核心缺陷:虚假关联、知识矛盾、泛化受限。
- 适合安全研究员和大模型应用开发者参考。
大型语言模型(LLMs)被广泛用于辅助安全分析师应对快速演变的网络威胁,提供威胁情报以支持漏洞评估和事件响应。尽管近期研究证明LLM可胜任多种威胁情报任务,如威胁分析、漏洞检测与入侵防御,但实际部署中仍存在显著性能差距。本文探究了LLM在威胁情报中的内在脆弱性,聚焦于威胁环境本身带来的挑战,而非模型架构问题。通过在多个威胁情报基准和真实威胁报告上进行大规模评估,提出一种融合分层、自回归优化与人工监督的新分类方法,可靠识别失败案例。经大量实验与人工核查,揭示三类根本性缺陷:虚假相关性、矛盾知识和泛化能力受限,限制了LLM在有效支持威胁情报中的表现。据此,为设计更鲁棒的LLM驱动威胁情报系统提供了可操作洞见,推动未来研究发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are intensively used to assist security analysts in counteracting the rapid exploitation of cyber threats, wherein LLMs offer cyber threat intelligence (CTI) to support vulnerability assessment and incident response. While recent work has shown that LLMs can support a wide range of CTI tasks such as threat analysis, vulnerability detection, and intrusion defense, significant performance gaps persist in practical deployments. In this paper, we investigate the intrinsic vulnerabilities of LLMs in CTI, focusing on challenges that arise from the nature of the threat landscape itself rather than the model architecture. Using large-scale evaluations across multiple CTI benchmarks and real-world threat reports, we introduce a novel categorization methodology that integrates stratification, autoregressive refinement, and human-in-the-loop supervision to reliably analyze failure instances. Through extensive experiments and human inspections, we reveal three fundamental vulnerabilities: spurious correlations, contradictory knowledge, and constrained generalization, that limit LLMs in effectively supporting CTI. Subsequently, we provide actionable insights for designing more robust LLM-powered CTI systems to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。