用推理模型提升数字取证结果可解释性,验证其实际效果
Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
- 引入gpt-oss模型,通过内部推理过程增强输出可解释性
- 中等推理层级下能有效辅助解释与验证结果,高阶推理无明显提升
- 适合关注模型可信度的取证人员或法律合规团队
大语言模型在数字取证中的应用已广泛研究,虽通过微调优化了性能,但结果可解释性不足限制了其实际与法律可用性。近期出现的推理型语言模型通过内部推理机制处理逻辑任务,但用户通常仅见最终答案,难见推理过程。本文首次评估推理模型在数字取证中的潜力,采用gpt-oss(可本地部署)进行四类典型用例测试,结合定量指标与定性分析。结果表明,在中等推理层级下,该机制有助于解释和验证模型输出,但支持有限;更高推理层级并未显著提升响应质量。
原文摘要 · Abstract (English)
The use of large language models in digital forensics has been widely explored. Beyond identifying potential applications, research has also focused on optimizing model performance for forensic tasks through fine-tuning. However, limited result explainability reduces their operational and legal usability. Recently, a new class of reasoning language models has emerged, designed to handle logic-based tasks through an `internal reasoning' mechanism. Yet, users typically see only the final answer, not the underlying reasoning. One of these reasoning models is gpt-oss, which can be deployed locally, providing full access to its underlying reasoning process. This article presents the first investigation into the potential of reasoning language models for digital forensics. Four test use cases are examined to assess the usability of the reasoning component in supporting result explainability. The evaluation combines a new quantitative metric with qualitative analysis. Findings show that the reasoning component aids in explaining and validating language model outputs in digital forensics at medium reasoning levels, but this support is often limited, and higher reasoning levels do not enhance response quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。