arXiv:2602.24047cs.NIcs.CR2026-02

无监督聚类+增量更新,动态识别物联网设备流量特征。

Unsupervised Baseline Clustering and Incremental Adaptation for IoT Device Traffic Profiling

  • 用DBSCAN无监督聚类提取设备流量基线,匹配真实标签效果好。
  • 增量更新中BIRCH每步仅需0.13秒,新设备聚类纯度达0.87。
  • 兼顾效率与适应性,适合持续演化的物联网安全场景。

物联网设备数量增长与异构性带来安全挑战,静态识别模型易因流量演变而退化。本文提出两阶段、基于流特征的无监督物联网设备流量画像与增量更新方法,基于德肯物联网数据集的长期捕获数据进行评估。在基线画像阶段,密度聚类(DBSCAN)有效分离异常数据,相比其他经典方法在匹配真实设备标签上表现最佳(NMI 0.78),且在聚类纯度上优于中心点聚类。在增量适应阶段,评估流式聚类方法发现,BIRCH支持高效更新(每次0.13秒),对未见新设备形成较一致的聚类(纯度0.87),但对新流量覆盖有限(占比0.72),且适应后已知设备准确率略有下降(0.71)。整体结果揭示了高纯度静态画像与增量聚类灵活性之间的实际权衡。

原文摘要 · Abstract (English)

The growth and heterogeneity of IoT devices create security challenges where static identification models can degrade as traffic evolves. This paper presents a two-stage, flow-feature-based pipeline for unsupervised IoT device traffic profiling and incremental model updating, evaluated on selected long-duration captures from the Deakin IoT dataset. For baseline profiling, density-based clustering (DBSCAN) isolates a substantial outlier portion of the data and produces the strongest alignment with ground-truth device labels among tested classical methods (NMI 0.78), outperforming centroid-based clustering on cluster purity. For incremental adaptation, we evaluate stream-oriented clustering approaches and find that BIRCH supports efficient updates (0.13 seconds per update) and forms comparatively coherent clusters for a held-out novel device (purity 0.87), but with limited capture of novel traffic (share 0.72) and a measurable trade-off in known-device accuracy after adaptation (0.71). Overall, the results highlight a practical trade-off between high-purity static profiling and the flexibility of incremental clustering for evolving IoT environments.

物联网无监督学习聚类增量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。