首个针对日志型入侵检测的LLM评测基准,揭示模型在复杂场景下表现大幅下降
HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection

- 构建统一日志数据集并转换为LLM可处理格式
- 模型在噪声大、复杂的日志中准确率骤降,MCC常低于0.5
- 发现模型存在过度敏感或过于保守两种典型行为模式
现有基准已推进大语言模型在渗透测试与漏洞识别等网络安全任务中的评估,但主机级入侵检测(HIDS)这一关键任务仍缺乏系统评测。本文提出HIDBench,用于评估大语言模型在基于系统日志的入侵检测中的能力。该任务需对大规模、嘈杂且高度不平衡的日志进行细粒度推理,恶意与正常行为的复杂交互使可靠检测极具挑战。本基准整合DARPA-E3、DARPA-E5和NodLink三个公开日志数据集,设计数据构建流水线,将原始主机遥测数据转化为适合LLM输入的格式,在真实入侵检测场景下实现系统性评估。对前沿模型的评测显示显著性能差距:多数模型在简单数据集上精度高于0.8,但在日志更嘈杂复杂时,马修斯相关系数(MCC)频繁低于0.5,误报率急剧上升。进一步分析揭示两种典型模型行为:低误报的保守型检测器与高告警量的过敏感模型。结果表明,尽管LLM在HIDS中潜力巨大,其效果高度依赖数据复杂度,可靠部署需精细系统设计。
原文摘要 · Abstract (English)
Recent benchmark efforts have advanced the evaluation of large language models (LLMs) in cybersecurity, including tasks such as penetration testing and vulnerability identification. However, a critical cybersecurity task, namely intrusion detection from system logs, remains unexplored. In this work, we present a new benchmark to assess LLMs' capabilities in supporting host-based intrusion detection systems (HIDS). This task requires fine-grained reasoning over large-scale, noisy, and highly imbalanced system logs, where complex interactions between benign and malicious activities make reliable detection challenging. Our benchmark unifies three public system log datasets, DARPA-E3, DARPA-E5, and NodLink, and introduces a data construction pipeline that transforms raw host telemetry into LLM-compatible inputs, enabling systematic evaluation under realistic intrusion detection settings. Our evaluation of frontier LLMs reveals substantial performance gaps across datasets. While many models achieve high precision (often above 0.8) on simpler datasets, their performance degrades significantly as system logs become noisier and more complex, with MCC frequently dropping below 0.5 and false positive rates increasing sharply. We further analyze model behavior and identify distinct regimes, including conservative detectors with low false positive rates and over-sensitive models that generate excessive alerts. Overall, our results highlight that while LLMs show strong potential for HIDS, their effectiveness is highly sensitive to data complexity, and robust system design is essential for reliable deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。