构建5000+条真实恶意软件分析问答对,助力数字取证研究
ForensicsData: A Digital Forensics Dataset for Large Language Models
- 从真实报告提取结构化数据,用大模型生成问答对
- 超5000个三元组,经专业评估确保术语准确性
- 适合取证研究者、安全工具开发者使用
日益复杂的网络事件给数字取证调查带来挑战,尤其在证据收集与分析方面。由于伦理、法律和隐私顾虑,公开可用资源仍十分有限,而真实数据集对研究与工具开发至关重要。为填补这一空白,我们提出ForensicsData,一个源自真实恶意软件分析报告的大型问答-上下文-答案(Q-C-A)数据集,包含超过5,000个三元组。通过独特的工作流:提取结构化数据,利用大语言模型(LLMs)转换为Q-C-A格式,并采用专项评估流程验证质量。在多个模型中,Gemini 2 Flash在内容与取证术语对齐上表现最佳。ForensicsData旨在推动数字取证研究,支持可复现实验并促进学术协作。
原文摘要 · Abstract (English)
The growing complexity of cyber incidents presents significant challenges for digital forensic investigators, especially in evidence collection and analysis. Public resources are still limited because of ethical, legal, and privacy concerns, even though realistic datasets are necessary to support research and tool developments. To address this gap, we introduce ForensicsData, an extensive Question-Context-Answer (Q-C-A) dataset sourced from actual malware analysis reports. It consists of more than 5,000 Q-C-A triplets. A unique workflow was used to create the dataset, which extracts structured data, uses large language models (LLMs) to transform it into Q-C-A format, and then uses a specialized evaluation process to confirm its quality. Among the models evaluated, Gemini 2 Flash demonstrated the best performance in aligning generated content with forensic terminology. ForensicsData aims to advance digital forensics by enabling reproducible experiments and fostering collaboration within the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。