arXiv:2512.03070cs.LGcs.AI2025-12综述被引 1

提出基于预拓扑空间的混合数据聚类方法,解决异构数据聚类难题。

Mixed Data Clustering Survey and Challenges

  • 基于预拓扑空间设计新型聚类算法,支持异构数据融合。
  • 在真实数据集上优于经典数值聚类与现有预拓扑方法。
  • 适合需要可解释性结果的工业级大数据分析场景。

大数据时代的到来彻底改变了产业界对信息的管理与分析方式,带来了前所未有的数据规模、速度与多样性。在此背景下,混合数据聚类成为关键挑战,亟需能够有效利用异质数据类型(包括数值型与类别型变量)的新方法。传统聚类技术多针对同质数据集设计,难以应对混合数据带来的额外复杂性,凸显了专为该场景定制方法的必要性。层次化与可解释算法在此尤为关键,因其能提供结构清晰、可解释的聚类结果,支持决策制定。本文提出一种基于预拓扑空间的聚类方法,并通过与经典数值聚类算法及现有预拓扑方法的对比,评估其在大数据范式下的性能与有效性。

原文摘要 · Abstract (English)

The advent of the big data paradigm has transformed how industries manage and analyze information, ushering in an era of unprecedented data volume, velocity, and variety. Within this landscape, mixed-data clustering has become a critical challenge, requiring innovative methods that can effectively exploit heterogeneous data types, including numerical and categorical variables. Traditional clustering techniques, typically designed for homogeneous datasets, often struggle to capture the additional complexity introduced by mixed data, underscoring the need for approaches specifically tailored to this setting. Hierarchical and explainable algorithms are particularly valuable in this context, as they provide structured, interpretable clustering results that support informed decision-making. This paper introduces a clustering method grounded in pretopological spaces. In addition, benchmarking against classical numerical clustering algorithms and existing pretopological approaches yields insights into the performance and effectiveness of the proposed method within the big data paradigm.

混合数据聚类可解释性预拓扑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。