提出可跨网络环境通用的物联网设备识别方法,解决现有模型泛化能力差的问题。
GeMID: Generalizable Models for IoT Device Identification
- 用遗传算法结合多环境数据优化特征与模型选择,提升泛化性
- 实验证明传统滑动窗口和流量统计方法泛化能力差,统计方法不可靠
- 适合关注物联网安全、模型鲁棒性的研究人员与工程师
随着物联网设备激增,设备识别(DI)成为保障安全的关键环节,通过分析流量模式区分设备并发现脆弱设备。然而,现有基于机器学习的识别方法常忽视跨不同网络环境的泛化能力问题。本文提出一种新框架,通过两阶段流程应对该挑战:首先,利用遗传算法结合外部反馈及多环境数据,优化特征与模型选择以增强泛化性;其次,在独立新数据集上测试模型,严格评估其泛化表现。实验表明,主流方法如滑动窗口和流统计因依赖网络特定特征而泛化性差,且广泛使用的统计方法因受网络特性影响而非设备固有属性,导致大量现有研究有效性存疑。本研究推动了物联网安全与设备识别领域的发展,为提升模型效能与降低网络风险提供新思路。
原文摘要 · Abstract (English)
With the proliferation of devices on the Internet of Things (IoT), ensuring their security has become paramount. Device identification (DI), which distinguishes IoT devices based on their traffic patterns, plays a crucial role in both differentiating devices and identifying vulnerable ones, closing a serious security gap. However, existing approaches to DI that build machine learning models often overlook the challenge of model generalizability across diverse network environments. In this study, we propose a novel framework to address this limitation and to evaluate the generalizability of DI models across data sets collected within different network environments. Our approach involves a two-step process: first, we develop a feature and model selection method that is more robust to generalization issues by using a genetic algorithm with external feedback and datasets from distinct environments to refine the selections. Second, the resulting DI models are then tested on further independent datasets to robustly assess their generalizability. We demonstrate the effectiveness of our method by empirically comparing it to alternatives, highlighting how fundamental limitations of commonly employed techniques such as sliding window and flow statistics limit their generalizability. Moreover, we show that statistical methods, widely used in the literature, are unreliable for device identification due to their dependence on network-specific characteristics rather than device-intrinsic properties, challenging the validity of a significant portion of existing research. Our findings advance research in IoT security and device identification, offering insight into improving model effectiveness and mitigating risks in IoT networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。