arXiv:2608.11492cs.CRcs.LG2026-08

构建首个人工验证的物联网固件漏洞检测基准,提升跨平台泛化能力。

Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware

论文配图:Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware
图 1 · 摘自论文原文
  • 基于人工验证的GitHub数据构建跨语料库评测基准IoTVulBench。
  • 最优模型在0.5%误报率下仅漏检21%漏洞,较传统分析器提升0.42 MCC。
  • 课程学习与集成策略显著提升性能,适合安全研究与工业部署场景。

物联网固件漏洞检测受限于生态异构性、资源受限平台及基准质量不足。现有数据集多为合成或通用数据,缺乏人工验证与污染筛查标注,导致跨语料库泛化在训练来源、架构和课程设计上研究不足。本研究提出IoTVulBench,一个经三位专家评审的人工验证基准,用于跨语料库固件漏洞检测。该基准在五个架构、两种调优方法、三种课程策略下进行污染筛选的保留测试,并开展集成、蒸馏与鲁棒性分析。在欠采样匹配条件下,使用IoTVulBench训练的模型达到最高MCC值0.58,优于PrimeVul(0.44)和D2A(0.39)。分阶段课程学习使MCC提升至0.69,多样性优化集成达0.73。相比最强对比基线(静态分析器,MCC=0.31),MCC提升0.42;相比最强单源数据集PrimeVul,提升0.29。在0.5%误报率下,模型仅漏检21%漏洞,而对比基线为71%。模型在标识符重命名下仍保持86%性能,校准良好。结果表明,领域匹配数据与课程设计是推动固件漏洞检测泛化的关键,而非模型规模,同时提供可部署的基准配置。

原文摘要 · Abstract (English)

IoT firmware vulnerability detection is constrained by ecosystem heterogeneity, resource-limited platforms, and benchmark quality limitations. Existing datasets are often synthetic or general-purpose and lack human-verified, contamination-screened annotations, leaving cross-corpus generalization across training sources, architectures, and curriculum design underexplored. In this study, we have introduced IoTVulBench, a human-verified benchmark for cross-corpus firmware vulnerability detection. IoTVulBench was built from GitHub repositories, validated by three expert reviewers, and evaluated on a contamination-screened held-out target across five architectures, two tuning methods, and three curriculum strategies, with ensemble, distillation, and robustness analyses. Models trained on IoTVulBench reached the highest Matthews Correlation Coefficient (MCC) among undersampling-matched single-source datasets, at 0.58 versus 0.44 for PrimeVul and 0.39 for D2A. Staged curriculum learning raised MCC to 0.69, and a diversity-optimized ensemble reached 0.73. This gain represents a 0.42 MCC improvement over the strongest reference comparator, a static analyzer with an MCC of 0.31, and a 0.29 MCC improvement over the strongest single-source dataset, PrimeVul. At a 0.5% false-positive rate, the model missed only 21% of vulnerabilities versus 71% for the comparator. The model also retained 86% of its performance under identifier renaming, with strong calibration. These results indicate that domain-matched training data and curriculum design, rather than model scale alone, drive generalization in firmware vulnerability detection, and yield both a benchmark and deployment-ready configurations for IoT security.

漏洞检测物联网安全基准测试课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。