arXiv:2503.05772cs.LG2025-03被引 1

用最小生成树和最短路径构建网络特征,实现复杂模式的高效分类。

Complex Networks for Pattern-Based Data Classification

  • 基于最小生成树与单源最短路径构造网络度量,捕捉数据内在模式。
  • 在合成与真实数据集上表现优于传统方法,准确率显著提升。
  • 仅需单一网络度量,免去参数调优,模型更简洁且泛化性强。

数据分类技术将数据或特征空间划分为对应特定类别的子空间,通常依赖距离、分布等物理特征。然而,对于嵌入在数据中的复杂模式,该方法面临挑战。复杂网络能有效捕获内部关系与类别结构,适用于高层次分类。尽管已有多种基于复杂网络的分类方法,但利用模式形成进行高层次分类尚未被充分探索。本文提出两种基于网络的分类技术,采用从最小生成树和单源最短路径中提取的独特度量,评估由每类数据自身构成的内在网络所呈现的数据模式。我们在多个合成与真实世界数据集上验证了所提方法,相比现有经典及机器学习分类方法,取得了有竞争力的数值结果。此外,所提模型相较以往高层分类技术具备以下优势:(1) 引入单一网络度量表征数据模式,无需调节度量间权重,显著简化模型同时提升分类效果;(2) 提出度量敏感性强,分类性能优异。实验表明,该方法在复杂模式识别任务中具有显著潜力。

原文摘要 · Abstract (English)

Data classification techniques partition the data or feature space into smaller sub-spaces, each corresponding to a specific class. To classify into subspaces, physical features e.g., distance and distributions are utilized. This approach is challenging for the characterization of complex patterns that are embedded in the dataset. However, complex networks remain a powerful technique for capturing internal relationships and class structures, enabling High-Level Classification. Although several complex network-based classification techniques have been proposed, high-level classification by leveraging pattern formation to classify data has not been utilized. In this work, we present two network-based classification techniques utilizing unique measures derived from the Minimum Spanning Tree and Single Source Shortest Path. These network measures are evaluated from the data patterns represented by the inherent network constructed from each class. We have applied our proposed techniques to several data classification scenarios including synthetic and real-world datasets. Compared to the existing classic high-level and machine-learning classification techniques, we have observed promising numerical results for our proposed approaches. Furthermore, the proposed models demonstrate the following distinguished features in comparison to the previous high-level classification techniques: (1) A single network measure is introduced to characterize the data pattern, eliminating the need to determine weight parameters among network measures. Therefore, the model is largely simplified, while obtaining better classification results. (2) The metrics proposed are sensitive and used for classification with competitive results.

复杂网络模式分类数据挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。