通过联合嵌入空间提升代谢物谱图注释准确率
JESTR: Joint Embedding Space Technique for Ranking Candidate Molecules for the Annotation of Untargeted Metabolomics Data
- 将分子与谱图映射到同一嵌入空间,用相似度排序候选结构
- 在3个数据集上,前5名排名准确率比现有工具高23.6%至71.6%
- 训练时加入候选分子正则化,显著提升区分能力,适合代谢组学研究者
代谢组学中注释的核心挑战是为质谱碎片模式分配分子结构。尽管分子到谱图及谱图到分子指纹(FP)预测已有进展,注释率仍较低。本文提出一种新范式JESTR:利用分子与其对应谱图是同一数据的多视角这一洞察,将两者表示嵌入到联合空间中,通过查询谱图与候选结构嵌入间的余弦相似度进行排序。我们在三个数据集上评估了JESTR,相较于mol-to-spec和spec-to-FP工具,平均在rank@[1-5]上提升23.6%至71.6%。训练时引入候选分子正则化,使rank@1性能提升11.4%,增强模型区分目标与候选分子的能力。在MassSpecGym基准数据集子集上,相较公开预训练的SIRIUS和CFM-ID模型,JESTR分别提升31%和238%。JESTR为实现精准注释提供了新路径,有助于揭示代谢组学深层信息。
原文摘要 · Abstract (English)
Motivation: A major challenge in metabolomics is annotation: assigning molecular structures to mass spectral fragmentation patterns. Despite recent advances in molecule-to-spectra and in spectra-to-molecular fingerprint prediction (FP), annotation rates remain low. Results: We introduce in this paper a novel paradigm (JESTR) for annotation. Unlike prior approaches that explicitly construct molecular fingerprints or spectra, JESTR leverages the insight that molecules and their corresponding spectra are views of the same data and effectively embeds their representations in a joint space. Candidate structures are ranked based on cosine similarity between the embeddings of query spectrum and each candidate. We evaluate JESTR against mol-to-spec and spec-to-FP annotation tools on three datasets. On average, for rank@[1-5], JESTR outperforms other tools by 23.6%-71.6%. We further demonstrate the strong value of regularization with candidate molecules during training, boosting rank@1 performance by 11.4% and enhancing the model's ability to discern between target and candidate molecules. When comparing JESTR's performance against that of publicly available pretrained models of SIRIUS and CFM-ID on appropriate subsets of MassSpecGym benchmark dataset, JESTR outperforms these tools by 31% and 238%, respectively. Through JESTR, we offer a novel promising avenue towards accurate annotation, therefore unlocking valuable insights into the metabolome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。