提出可评估风险的恶意软件分析框架,自动处理93%样本,误接受率仅0.136%
EGAMA-RC: Risk-Calibrated Evidence-Gated Adaptive Malware Analysis for Robust and Interpretable Memory-Forensic Triage
- 基于特征优化与证据门控,动态分流高风险、不确定或新型样本
- 93.12%样本自动通过,准确率达99.86%,误接受率0.136%
- 兼顾可解释性与效率,适合需要低误判的实战安全研判
机器学习恶意软件检测器在干净数据上表现优异,但实际取证需兼顾不确定性、新颖性、鲁棒性、可解释性、延迟和人工审查成本。本文提出EGAMA-RC,一种面向内存取证的鲁棒且可解释的恶意软件分类框架。该框架结合SHAP引导的特征精炼、数据集特异性优化、模型池评估、对抗测试与开放家族测试、新颖性评分、解释条件证据及运行时感知路由。低风险样本自动通过,而高风险、不确定或潜在新型样本则被转至人工审查、升级或新颖性敏感处理。在三个恶意软件数据集与冻结多种子协议下,混合门控机制接收93.12%的样本,接受准确率达99.86%,误接受率为0.136%。新颖性校准有效缓解过度审查行为,同时保持低安全隐患。XGBoost实现轻量快速路径推理,单样本中位/95%延迟为0.0054/0.0059毫秒。结果表明,可靠恶意软件分析需风险校准路由、新颖性感知与受控人工干预,而非仅依赖分类准确率。
原文摘要 · Abstract (English)
Machine-learning malware detectors often achieve high clean-data accuracy, but operational triage also requires evidence about uncertainty, novelty, robustness, interpretability, latency, and review cost. This paper presents EGAMA-RC, a risk-calibrated evidence-gated framework for memory-forensic malware triage. Building on SHAP-guided feature refinement, EGAMA-RC combines dataset-specific refinement, model-pool evaluation, adversarial and open-family testing, novelty scoring, explanation-conditioned evidence, and runtime-aware routing. Low-risk samples are accepted automatically, while uncertain, high-risk, or potentially novel cases are routed to review, escalation, or novelty-aware handling. Across three malware datasets and a frozen multi-seed protocol, the selected hybrid gate accepts 93.12% of pooled samples with 99.86% accepted accuracy and a 0.136% false-accept rate. Novelty calibration reduces over-restrictive review behavior while preserving a low unsafe-accept profile. XGBoost provides lightweight fast-path inference with p50/p95 latency of 0.0054/0.0059 ms per sample. The results show that dependable malware analysis requires risk-calibrated routing, novelty awareness, and controlled analyst review, not classification accuracy alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。