arXiv:2603.20829cs.LG2026-03被引 1

提出统一框架与工业视角,解决图聚类方法在真实场景中的落地难题。

Beyond the Academic Monoculture: A Unified Framework and Industrial Perspective for Attributed Graph Clustering

  • 构建编码-聚类-优化三模块框架,便于方法组合与对比
  • 指出学术评估依赖小规模同质数据,忽视效率与真实场景
  • 面向工业部署提出可扩展、抗异质性的实用策略

属性图聚类(AGC)是一种基础的无监督任务,通过联合建模结构拓扑和节点属性来划分节点为紧密群体。尽管图神经网络和自监督学习推动了大量AGC方法的发展,但学术基准性能与真实工业部署的严苛需求之间仍存在巨大鸿沟。本文从三个互补视角进行系统综述:首先提出编码-聚类-优化分类框架,将多样算法分解为三个正交可组合模块;其次批判现有评估协议,揭示对小型同质引用网络的过度依赖、仅使用监督指标评价无监督任务、以及计算可扩展性长期被忽视的问题,倡导融合语义对齐、结构完整性和效率分析的综合评估标准;最后结合大规模、严重异质性和表格式特征噪声等实际约束,基于配套基准的实证分析,提出可行工程策略,并规划未来研究方向,优先发展鲁棒异质性编码器、可扩展联合优化及无监督模型选择标准以满足生产级要求。

原文摘要 · Abstract (English)

Attributed Graph Clustering (AGC) is a fundamental unsupervised task that partitions nodes into cohesive groups by jointly modeling structural topology and node attributes. While the advent of graph neural networks and self-supervised learning has catalyzed a proliferation of AGC methodologies, a widening chasm persists between academic benchmark performance and the stringent demands of real-world industrial deployment. To bridge this gap, this survey provides a comprehensive, industrially grounded review of AGC from three complementary perspectives. First, we introduce the Encode-Cluster-Optimize taxonomic framework, which decomposes the diverse algorithmic landscape into three orthogonal, composable modules: representation encoding, cluster projection, and optimization strategy. This unified paradigm enables principled architectural comparisons and inspires novel methodological combinations. Second, we critically examine prevailing evaluation protocols to expose the field's academic monoculture: a pervasive over-reliance on small, homophilous citation networks, the inadequacy of supervised-only metrics for an inherently unsupervised task, and the chronic neglect of computational scalability. In response, we advocate for a holistic evaluation standard that integrates supervised semantic alignment, unsupervised structural integrity, and rigorous efficiency profiling. Third, we explicitly confront the practical realities of industrial deployment. By analyzing operational constraints such as massive scale, severe heterophily, and tabular feature noise alongside extensive empirical evidence from our companion benchmark, we outline actionable engineering strategies. Furthermore, we chart a clear roadmap for future research, prioritizing heterophily-robust encoders, scalable joint optimization, and unsupervised model selection criteria to meet production-grade requirements.

图聚类工业落地评估标准异质性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。