比较多项式树聚类方法,发现归一化距离效果最佳
A Comparison of Polynomial-Based Tree Clustering Methods
- 用树区分多项式编码树结构,支持高效聚类
- 基于归一化距离的聚类方法准确率最高
- 适合生物树结构数据分析与机器学习应用
树结构广泛存在于生命科学领域,如系统发育、发育生物学和核酸结构。树可表示非编码RNA的二级结构,直接关联其功能。近年来测序技术和人工智能的发展产生了大量可表示为树结构的生物数据,亟需新的树结构数据分析方法。树多项式提供了一种计算高效、可解释且全面的矩阵编码方式,兼容主流数据分析工具。基于树多项式之间坎贝尔距离的机器学习方法已被用于分析系统发育和核酸结构。本文比较了基于树区分多项式的不同距离在树聚类中的性能,并实现了两种基础自编码器模型用于树聚类。结果表明,采用基线归一化距离的方法在所有对比方法中具有最高的聚类准确率。
原文摘要 · Abstract (English)
Tree structures appear in many fields of the life sciences, including phylogenetics, developmental biology and nucleic acid structures. Trees can be used to represent RNA secondary structures, which directly relate to the function of non-coding RNAs. Recent developments in sequencing technology and artificial intelligence have yielded numerous biological data that can be represented with tree structures. This requires novel methods for tree structure data analytics. Tree polynomials provide a computationally efficient, interpretable and comprehensive way to encode tree structures as matrices, which are compatible with most data analytics tools. Machine learning methods based on the Canberra distance between tree polynomials have been introduced to analyze phylogenies and nucleic acid structures. In this paper, we compare the performance of different distances in tree clustering methods based on a tree distinguishing polynomial. We also implement two basic autoencoder models for clustering trees using the polynomial. We find that the distance based methods with entry-level normalized distances have the highest clustering accuracy among the compared methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。