arXiv:2601.09029cs.CRcs.AI2026-01

用大模型主动从网页情报中挖掘恶意线索,精准识别威胁。

Proactively Detecting Threats: A Novel Approach Using LLMs

  • 用大模型自动解析15个网页情报源提取威胁指标
  • Gemini 1.5 Pro识别恶意指标精度达0.958,召回率1.0
  • 首次系统评估大模型在主动威胁检测中的表现

企业安全面临日益复杂的恶意软件威胁,尤其在数字化运营扩展的背景下。本文首次系统评估大语言模型(LLMs)在主动识别非结构化网络威胁情报中指示性攻击特征(IOCs)的能力,区别于传统的被动恶意软件检测。我们构建了一个自动化系统,从15个基于网页的威胁报告源中提取指标,评估了六种LLM模型(Gemini、Qwen及Llama系列)。对包含2,658个指标的479个网页进行评估,其中含711个IPv4地址、502个IPv6地址和1,445个域名。结果显示性能差异显著:Gemini 1.5 Pro在恶意指标识别上达到0.958的精确率与0.788的特异性,对真实威胁实现100%召回率。

原文摘要 · Abstract (English)

Enterprise security faces escalating threats from sophisticated malware, compounded by expanding digital operations. This paper presents the first systematic evaluation of large language models (LLMs) to proactively identify indicators of compromise (IOCs) from unstructured web-based threat intelligence sources, distinguishing it from reactive malware detection approaches. We developed an automated system that pulls IOCs from 15 web-based threat report sources to evaluate six LLM models (Gemini, Qwen, and Llama variants). Our evaluation of 479 webpages containing 2,658 IOCs (711 IPv4 addresses, 502 IPv6 addresses, 1,445 domains) reveals significant performance variations. Gemini 1.5 Pro achieved 0.958 precision and 0.788 specificity for malicious IOC identification, while demonstrating perfect recall (1.0) for actual threats.

大模型威胁检测情报分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。