PIDS性能受评测方法影响大,简单白名单竟超复杂模型。
How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection
- 统一时间分离测试与验证集选择协议,避免过拟合
- 三个数据集上白名单性能优于多数学习模型
- 强调评测设计对结论的影响,适合安全研究者参考
基于溯源的入侵检测系统(PIDS)常报告优异性能,但其结论高度依赖于基准选择与评测协议。我们对满足审计、标注与校准要求的公开数据集重新评估代表性PIDS。聚焦经审计的DARPA TC E3数据集,采用统一协议——时间分离测试期、仅验证集选模型与阈值校准——检验架构主张的实证支持度。发现告警成功率与调查实用性可能严重背离:多个系统虽能触发攻击告警,却缺乏足够的进程级上下文支持取证分析。在四个主数据集中的三个上,仅基于训练时可执行文件名与路径构建的简单白名单,在关键操作点指标上达到或超越选定的学习基线,表明部分系统性能反映的是词汇新颖性而非深层溯源建模能力。通过特征完备性与字段熵量化语义信号质量,解释为何多个审计后的E3数据集支持告警性能但无法可靠区分模型架构;而Theia结合最丰富的语义信号,在参考模型下实现更清晰的排序提升与节点级恢复。总体表明,解读PIDS架构主张必须结合其产生结论的基准属性与评测协议。
原文摘要 · Abstract (English)
Provenance-based intrusion detection systems (PIDS) frequently report strong performance, but the conclusions drawn from these results can be highly sensitive to benchmarking choices and evaluation protocols. We investigate this dependency by re-evaluating representative PIDS on public datasets that meet our audit, labeling, and calibration requirements. Focusing primarily on the audited DARPA TC E3 datasets, we apply a unified protocol with temporally separated test periods and validation-only checkpoint selection and threshold calibration, and ask which architectural claims are empirically supported. We find that alerting success and investigation utility can diverge sharply, as several systems surface attacks without providing enough process-level context to support forensic investigation. On three of the four primary datasets, a simple allowlist built from training executable names and paths matches or exceeds the selected learned baselines on key operating-point metrics, suggesting that much of these systems' measured performance reflects lexical novelty rather than richer provenance modeling. Quantifying semantic signal quality through feature completeness and field entropy helps explain why several audited E3 datasets support alerting performance without reliably separating model architectures, whereas Theia combines the richest semantic signal with the clearest improvements in ranking and node-level recovery by our reference model. Overall, these findings reinforce the importance of interpreting architectural claims in PIDS together with the benchmark properties and evaluation protocol that produced them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。