用可解释的塔斯林机检测恶意PDF,准确率达98.02%
Leveraging Interpretable Tsetlin Machine for PDF Malware Detection

- 通过静态分析提取PDF特征,用规则学习分类
- 在RIT-PDFMal-2026数据集上准确率达98.02%
- 结果可解释,适合需要透明决策的场景
在数字时代,便携文档格式(PDF)因其平台无关性和丰富功能,成为存储与交换数字文档的主流格式。然而,这些特性也使其成为网络攻击者的首选载体,攻击者常在看似合法的文档中嵌入恶意代码以破坏目标系统。本文提出一种基于可解释塔斯林机(Tsetlin Machine, TM)的PDF恶意软件检测新框架。该框架通过静态分析提取PDF文档的关键特征,无需执行文件即可进行分类,利用规则学习实现对良性与恶意PDF文档的精准区分。在RIT-PDFMal-2026数据集上的数值评估表明,该框架准确率达到98.02%,优于多种先进机器学习分类器。此外,框架具备内在可解释性,能透明展示分类依据。在树莓派上实现边缘部署,支持实时、设备端的PDF恶意软件检测。结合更高准确率、计算效率与内在可解释性,该框架为实际应用中的PDF恶意软件检测提供了有力解决方案。
原文摘要 · Abstract (English)
In the digital era, Portable Document Format (PDF) is one of the most widely used file formats for storing and exchanging digital documents due to its platform independence and rich functionality. However, these same capabilities have also made PDF files an attractive attack vector for cyberattackers, who embed malicious code within seemingly legitimate documents to compromise target systems. This paper presents a novel interpretable Tsetlin Machine (TM)-based framework for PDF malware detection. The proposed framework extracts salient features from PDF documents through static analysis without executing the files and employs rule-based learning to accurately classify benign and malicious PDF documents. Numerical evaluation on the RIT-PDFMal-2026 dataset demonstrates that the proposed framework achieves an accuracy of 98.02%, outperforming several state-of-the-art machine learning classifiers. Moreover, the proposed framework provides intrinsic interpretability by transparently explaining its classification decisions. Edge deployment on a Raspberry Pi further supports real-time, on-device PDF malware detection. The combination of better accuracy, computational efficiency, and intrinsic interpretability makes the proposed framework a promising solution for practical PDF malware detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。