arXiv:2409.14796cs.LGcs.AI2024-09被引 12

用无监督方法检测动态数据流异常,对不平衡数据更有效。

Research on Dynamic Data Flow Anomaly Detection based on Machine Learning

  • 通过多维特征提取与聚类分析,自动识别偏离正常流量的异常行为。
  • 在多种场景下检测准确率高,尤其在数据不平衡时表现稳定。
  • 适合缺乏标注数据的实时安全监控场景,如网络入侵检测。

当前网络攻击日益复杂多样,仅依赖代理、网关、防火墙和加密隧道已难以应对。因此,主动识别数据异常成为数据安全领域的研究热点。现有多数研究集中于样本均衡数据,导致在非均衡数据下的检测效果不佳。本文采用无监督学习方法检测动态数据流中的异常。首先从实时数据中提取多维特征,再利用聚类算法分析数据模式,从而自动识别潜在异常点。通过聚类相似数据,模型可在无需标签的情况下发现显著偏离正常流量的行为。实验结果表明,该方法在多种场景下均具备高检测准确率,尤其在非均衡数据条件下表现出强鲁棒性和良好适应性。

原文摘要 · Abstract (English)

The sophistication and diversity of contemporary cyberattacks have rendered the use of proxies, gateways, firewalls, and encrypted tunnels as a standalone defensive strategy inadequate. Consequently, the proactive identification of data anomalies has emerged as a prominent area of research within the field of data security. The majority of extant studies concentrate on sample equilibrium data, with the consequence that the detection effect is not optimal in the context of unbalanced data. In this study, the unsupervised learning method is employed to identify anomalies in dynamic data flows. Initially, multi-dimensional features are extracted from real-time data, and a clustering algorithm is utilised to analyse the patterns of the data. This enables the potential outliers to be automatically identified. By clustering similar data, the model is able to detect data behaviour that deviates significantly from normal traffic without the need for labelled data. The results of the experiments demonstrate that the proposed method exhibits high accuracy in the detection of anomalies across a range of scenarios. Notably, it demonstrates robust and adaptable performance, particularly in the context of unbalanced data.

异常检测无监督学习数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。