arXiv:2503.22448physics.opticscs.LG2025-03被引 3

比较三种聚类方法在流体透镜光学数据中的表现,发现混合方案更稳定可靠。

Comparison between neural network clustering, hierarchical clustering and k-means clustering: Applications using fluidic lenses

  • 用自组织映射分析15个泽尼克系数的分布结构,识别出8个关键聚类。
  • 层次聚类得cophenetic相关系数0.9651,确定7个最优簇;K均值聚类轮廓值达0.905。
  • 组合使用层次与K均值聚类比单一方法更稳定,避免参数变动导致结果漂移。

本文对比了神经网络聚类(NNC)、层次聚类(HC)和K均值聚类(KMC)在处理大规模数据集时的计算性能。针对流体透镜的波前传感器重建数据,采用15个泽尼克系数表征光相位畸变。通过自组织映射(SOM)分析权重距离、样本命中率、权重位置及平面分布,可视化系统结构特征。层次聚类结合差异性-链接矩阵计算,获得高一致性相关系数0.9651,设定不一致截断值0.8后确定7个簇用于系统分割。K均值聚类在划分5个非重叠簇时,轮廓平均值达到0.905,体现高效分割能力。相比之下,神经网络聚类表明15个变量可由8个核心簇协同描述。研究证实,结合层次聚类与K均值的混合策略比单独使用任一方法更可靠,因改变SOM大小或不一致截断值会引发聚类配置根本性变化。

原文摘要 · Abstract (English)

A comparison between neural network clustering (NNC), hierarchical clustering (HC) and K-means clustering (KMC) is performed to evaluate the computational superiority of these three machine learning (ML) techniques for organizing large datasets into clusters. For NNC, a self-organizing map (SOM) training was applied to a collection of wavefront sensor reconstructions, decomposed in terms of 15 Zernike coefficients, characterizing the optical aberrations of the phase front transmitted by fluidic lenses. In order to understand the distribution and structure of the 15 Zernike variables within an input space, SOM-neighboring weight distances, SOM-sample hits, SOM-weight positions and SOM-weight planes were analyzed to form a visual interpretation of the system's structural properties. In the case of HC, the data was partitioned using a combined dissimilarity-linkage matrix computation. The effectiveness of this method was confirmed by a high cophenetic correlation coefficient value (c=0.9651). Additionally, a maximum number of clusters was established by setting an inconsistency cutoff of 0.8, yielding a total of 7 clusters for system segmentation. In addition, a KMC approach was employed to establish a quantitative measure of clustering segmentation efficiency, obtaining a sillhoute average value of 0.905 for data segmentation into K=5 non-overlapping clusters. On the other hand, the NNC analysis revealed that the 15 variables could be characterized through the collective influence of 8 clusters. It was established that the formation of clusters through the combined linkage and dissimilarity algorithms of HC alongside KMC is a more dependable clustering solution than separate assessment via NNC or HC, where altering the SOM size or inconsistency cutoff can lead to completely new clustering configurations.

聚类分析光学建模流体透镜自组织映射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。