解决质谱大数据聚类的可扩展性与稳定性难题,实现高效精准分析。
A Flexible Adaptive Stable Clustering Algorithm for Archive-Scale Online Mass Spectrometry

- 提出动态系统框架,解耦相似度核与优化逻辑,提升稳定性。
- 在2500万条大气气溶胶质谱数据上实现线性运行时间,准确识别稀有工业痕迹。
- 适合环境科学、大数据分析领域研究者,用于处理海量质谱数据。
现代在线质谱技术产生海量多太字节数据流,对理解地球环境系统至关重要。但现有聚类方法在可扩展性、度量灵活性和算法稳定性之间存在权衡,制约了化学洞察的提取。本文提出柔性自适应稳定聚类(FASC),一种基于动力系统框架的方法,通过架构分离相似度核与严格优化逻辑,突破上述瓶颈。FASC采用密度增强的相似度选择规则与几何约束,确保确定性、顺序无关的收敛性。在标准机器学习基准上验证,聚类纯度超99.5%,调整兰德指数达0.99。将其部署于2500万条大气气溶胶质谱数据,实现严格线性经验运行时间(O(N)),自主揭示二次无机气溶胶的演化路径,并分离出丰度低于0.2%的超稀有工业示踪物,为环境大数据挖掘提供可扩展基础设施。
原文摘要 · Abstract (English)
Modern online mass spectrometry generates multi-terabyte data streams critical for understanding Earth's environmental systems. However, extracting actionable chemical insights from these repositories is impeded by a computational bottleneck: existing clustering methods force a compromise among scalability, metric flexibility, and algorithmic stability. Here, we introduce Flexible Adaptive Stable Clustering (FASC), a dynamical systems framework that resolves these constraints by architecturally decoupling the similarity kernel from rigorous optimization logic. Unlike legacy heuristics that suffer from stochastic drift and algorithmic blending, FASC employs a Density-Augmented Similarity Selection rule and geometric constraints to guarantee deterministic, order-independent convergence. After validating FASC on canonical machine-learning ground truths (achieving >99.5% cluster purity and 0.99 Adjusted Rand Index), we deployed the framework on 25 million mass spectra of atmospheric aerosols. Demonstrating strictly linear empirical runtime scaling (O(N)), FASC autonomously mapped atmospheric aging pathways of secondary inorganic aerosols while isolating ultra-rare industrial tracers (<0.2% abundance), providing a scalable infrastructure for mining environmental big data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。