arXiv:2505.19973cs.CRcs.AI2025-05被引 8

首个面向数字取证与响应的LLM评估基准,解决真实场景下的可信度问题。

DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response

  • 构建三类任务:认证题、实战取证挑战、真实磁盘/内存案例
  • 14个模型测试显示准确性差异大,引入新指标提升低准确场景评估能力
  • 适合安全研究者和取证工具开发者使用

数字取证与应急响应(DFIR)涉及分析数字证据以支持法律调查。大型语言模型(LLMs)在日志分析和内存取证等任务中展现潜力,但其易出错和幻觉问题在高风险场景中引发担忧。尽管关注度上升,目前尚无全面基准来评估LLMs在理论与实践双维度的表现。为此,我们提出DFIR-Metric,包含三个部分:(1) 知识评估:700道经专家审核的多选题,源自行业认证和官方文档;(2) 实用取证挑战:150个类似CTF的任务,测试多步推理与证据关联能力;(3) 实际分析:500个来自NIST计算机取证工具测试项目(CFTT)的真实磁盘与内存取证案例。我们对14个LLMs进行了评估,分析其准确率与试验间一致性,并引入新指标任务理解得分(TUS),用于更有效评估近零准确率场景下的模型表现。该基准为推动人工智能在数字取证领域的应用提供了严谨可复现的基础。所有脚本、数据及结果均公开于项目官网 https://github.com/DFIR-Metric。

原文摘要 · Abstract (English)

Digital Forensics and Incident Response (DFIR) involves analyzing digital evidence to support legal investigations. Large Language Models (LLMs) offer new opportunities in DFIR tasks such as log analysis and memory forensics, but their susceptibility to errors and hallucinations raises concerns in high-stakes contexts. Despite growing interest, there is no comprehensive benchmark to evaluate LLMs across both theoretical and practical DFIR domains. To address this gap, we present DFIR-Metric, a benchmark with three components: (1) Knowledge Assessment: a set of 700 expert-reviewed multiple-choice questions sourced from industry-standard certifications and official documentation; (2) Realistic Forensic Challenges: 150 CTF-style tasks testing multi-step reasoning and evidence correlation; and (3) Practical Analysis: 500 disk and memory forensics cases from the NIST Computer Forensics Tool Testing Program (CFTT). We evaluated 14 LLMs using DFIR-Metric, analyzing both their accuracy and consistency across trials. We also introduce a new metric, the Task Understanding Score (TUS), designed to more effectively evaluate models in scenarios where they achieve near-zero accuracy. This benchmark offers a rigorous, reproducible foundation for advancing AI in digital forensics. All scripts, artifacts, and results are available on the project website at https://github.com/DFIR-Metric.

数字取证LLM评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。