通过重排数据自动聚类,让相似项聚集、差异项分离。
Clustering data by reordering them
- 按元素间距离重排序列,利用局部相似性发现聚类结构。
- 在生物分子构象、基因序列等多类数据上验证有效,抗噪性强。
- 适合需要无监督分组的科研人员,尤其处理复杂高维数据。
在多个科学领域中,将元素分组以分别分析是标准流程。本文提出一种新算法,核心思想是同一类中的元素彼此相似,而与外部元素不同。通过根据元素间的距离对数据进行重排序,可自动完成分析,且参数直观易懂。算法显式考虑噪声,适用于数据驱动世界的多样化问题。已在生物分子构象、基因序列、细胞、图像及实验条件等场景中应用,表现良好。
原文摘要 · Abstract (English)
Grouping elements into families to analyse them separately is a standard analysis procedure in many areas of sciences. We propose herein a new algorithm based on the simple idea that members from a family look like each other, and don't resemble elements foreign to the family. After reordering the data according to the distance between elements, the analysis is automatically performed with easily-understandable parameters. Noise is explicitly taken into account to deal with the variety of problems of a data-driven world. We applied the algorithm to sort biomolecules conformations, gene sequences, cells, images, and experimental conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。