arXiv:2607.03653cs.CRcs.LG2026-07

融合卷积与视觉变压器,提升恶意软件图像分类准确率

ThreatVisionAI: A Hybrid CNN-ViT Framework for Image-Based Malware Classification

论文配图:ThreatVisionAI: A Hybrid CNN-ViT Framework for Image-Based Malware Classification
图 1 · 摘自论文原文
  • 用多尺度小波+CNN和ViT捕捉频率与全局关系特征
  • 在Malimg数据集上达98.01%准确率,少数类效果显著提升
  • 适合做图像化恶意软件分析的研究者与安全工程师

传统恶意软件检测方法难以泛化到混淆或未知威胁。本文提出ThreatVisionAI,一种混合框架,结合原始图像CNN、小波域CNN和视觉变换器(ViT),以捕获恶意软件图像中的互补空间、频域和全局依赖特征。小波域CNN提取多尺度频率信息,有助于区分关系密切的家族;ViT分支建模图像中长距离依赖关系。在Malimg数据集上,ThreatVisionAI达到98.01%的准确率和0.9742的加权F1分数,小波域特征对少数类及视觉相似家族带来可量化的性能提升。结果表明,频率感知与基于Transformer的表示能有效增强图像化恶意软件家族分类。

原文摘要 · Abstract (English)

Traditional malware detection methods struggle to generalize to obfuscated or previously unseen threats. This paper introduces ThreatVisionAI, a hybrid malware family classification framework that integrates a raw-image CNN, a wavelet-based CNN, and a Vision Transformer (ViT) to capture complementary spatial, frequency-domain, and global relational features in malware images. The wavelet-based CNN captures multi-scale frequency information that helps distinguish closely related families, while the ViT branch models long-range dependencies across the image. Evaluated on the Malimg dataset, ThreatVisionAI achieves 98.01% accuracy and a weighted F1 score of 0.9742, with wavelet-domain features providing measurable gains on minority and visually similar families. These results confirm that frequency-aware and transformer-based representations improve image-based malware family classification.

恶意软件分析图像分类ViT小波变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。