arXiv:2507.00039cs.LGcs.AI2025-07被引 1

对比38种图模式质量度量,发现常用方法效果不佳,提出预处理提升分类性能。

Pattern-Based Graph Classification: Comparison of Quality Measures and Importance of Preprocessing

  • 基于4个数学性质理论分析38种模式质量度量
  • 实验显示部分流行度量表现差,新聚类预处理可降维增效
  • 适合图分类、可解释性研究者参考

图分类旨在根据结构和属性特征对图进行分类,广泛应用于社交网络分析和生物信息学等领域。基于模式(即子图)的方法具有良好的可解释性,因分类所用模式可直接解读。为识别有意义的模式,需使用质量度量函数评估其判别能力。然而,文献中存在数十种此类度量,难以选择适用者。现有综述极少,且未聚焦图数据,导致普遍使用最流行的度量而缺乏验证。本文对38种文献中的质量度量进行比较分析,从四个数学性质进行理论刻画;利用公开数据集构建基准,并提出方法生成模式的黄金标准排序;在此基础上开展实证比较,涵盖模式排序与分类性能。此外,提出基于聚类的预处理步骤,将共现于相同图中的模式分组,以提升分类效果。实验表明该步骤能显著减少需处理的模式数,同时保持相当性能。结果还显示,部分文献中广泛使用的度量并未带来最优效果。

原文摘要 · Abstract (English)

Graph classification aims to categorize graphs based on their structural and attribute features, with applications in diverse fields such as social network analysis and bioinformatics. Among the methods proposed to solve this task, those relying on patterns (i.e. subgraphs) provide good explainability, as the patterns used for classification can be directly interpreted. To identify meaningful patterns, a standard approach is to use a quality measure, i.e. a function that evaluates the discriminative power of each pattern. However, the literature provides tens of such measures, making it difficult to select the most appropriate for a given application. Only a handful of surveys try to provide some insight by comparing these measures, and none of them specifically focuses on graphs. This typically results in the systematic use of the most widespread measures, without thorough evaluation. To address this issue, we present a comparative analysis of 38 quality measures from the literature. We characterize them theoretically, based on four mathematical properties. We leverage publicly available datasets to constitute a benchmark, and propose a method to elaborate a gold standard ranking of the patterns. We exploit these resources to perform an empirical comparison of the measures, both in terms of pattern ranking and classification performance. Moreover, we propose a clustering-based preprocessing step, which groups patterns appearing in the same graphs to enhance classification performance. Our experimental results demonstrate the effectiveness of this step, reducing the number of patterns to be processed while achieving comparable performance. Additionally, we show that some popular measures widely used in the literature are not associated with the best results.

图分类模式挖掘可解释性预处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。