arXiv:2606.05225q-bio.QMcs.LG2026-06

用自回归模型预测质谱中下一个出现的代谢物,提升未靶向检测的注释率。

The Language of Elution: Autoregressive Prediction of the Next Feature in Untargeted LC-HRMS Lipidomics

论文配图:The Language of Elution: Autoregressive Prediction of the Next Feature in Untargeted LC-HRMS Lipidomics
图 1 · 摘自论文原文
  • 将洗脱顺序视为语言序列,用LSTM和Transformer预测下一个质荷比区间。
  • 模型在四个临床队列上实现98.4%的准确率,误差仅3.6 Da。
  • 可跨仪器迁移,微调少量样本即可恢复性能,适合实际应用。

未靶向液相色谱-高分辨质谱(LC-HRMS)每样本可检测数千个分子特征,但仅有2%-20%获得可信结构注释。根本原因在于串联质谱(MS/MS)采集为反应式:仪器仅在离子出现后才选择前体,无法预知下一个洗脱物。本文将洗脱过程重新建模为自回归序列预测任务。由于反相洗脱顺序受疏水性控制,连续特征构成物理约束序列,类似语言中的词元。将质荷比(m/z)轴离散化为110个区间,利用长短期记忆(LSTM)与Transformer模型,基于五个无注释特征(m/z区间、质量缺陷、保留时间差、极性和强度排序)预测下一个洗脱的m/z区间。模型在4个临床脂质组学队列(共342份血浆样本;SCIEX TripleTOF 6600+,Waters CSH C18)上训练,LSTM达到98.4%的top-1准确率(top-5达99.99%,平均绝对误差3.6 Da),Transformer为98.0%。消融实验表明,自回归上下文贡献55.5个百分点,而任一单特征贡献不超过0.2个百分点:序列模式而非分子属性主导预测。模型在共享方法的独立Agilent 6530数据集上表现良好(r=0.999),但在不同柱化学(5.1% top-1)或极性模式下(2.6%)失败,证实其方法与模式依赖性。仅需2-5次质控注射微调,即可使丢失的准确率从2.6%恢复至近50%,说明跨条件部署只需极少校准。结果表明洗脱序列高度可预测,为预测性MS/MS采集提升未靶向代谢组学注释覆盖率奠定基础。

原文摘要 · Abstract (English)

Untargeted liquid chromatography-high-resolution mass spectrometry (LC-HRMS) detects thousands of molecular features per sample, yet only 2-20% receive confident structural annotations. A root cause of this "dark metabolome" is that tandem MS/MS acquisition is reactive: instruments select precursors only after ions appear, blind to what elutes next. We reframe chromatographic elution as an autoregressive sequence prediction task. Because reversed-phase elution order is governed by hydrophobicity, successive features form a physically constrained sequence, like tokens in language. We discretize the mass-to-charge (m/z) axis into 110 bins and train long short-term memory (LSTM) and Transformer models to predict the next eluting m/z bin from five annotation-free per-token features: m/z bin, mass defect, retention-time gap, polarity, and intensity rank. Trained on 15,242 features from four clinical lipidomics cohorts (342 plasma samples; SCIEX TripleTOF 6600+, Waters CSH C18), the LSTM reaches 98.4% top-1 accuracy (99.99% top-5; mean absolute error 3.6 Da) and the Transformer 98.0%. Ablation shows autoregressive context accounts for 55.5 percentage points while no single feature contributes more than 0.2 pp: the sequential pattern, not molecular properties, drives prediction. Models transfer across instruments sharing the method (r=0.999 on an independent Agilent 6530 dataset) but fail under a different column chemistry (5.1% top-1) or polarity mode (2.6%), confirming method- and mode-specificity. Fine-tuning on as few as two to five quality-control injections recovers held-out accuracy from 2.6% to nearly 50%, so cross-condition deployment needs minimal calibration. These results establish that elution sequences are highly predictable and lay the groundwork for predictive MS/MS acquisition to improve annotation coverage in untargeted metabolomics.

质谱分析自回归预测脂质组学机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。