构建统一分子质谱预测基准,解决模型评估不公与部署难问题
FlexMS: A Unified Public Benchmark for Molecule Tandem Mass Spectrum Prediction
- 设计模块化框架,统一不同数据源的分子编码与预测流程
- 引入难度感知诊断,支持在算力、数据量约束下的模型选型
- 支持真实代谢组学场景下的跨域评估,提升结果可复现性
串联质谱(MS/MS)是小分子鉴定的核心技术,但现有深度学习模型的谱图预测系统仍难以评估和部署。尽管新架构不断宣称达到最先进性能,但元数据条件不一致、预处理流程耦合,导致公平比较困难。此外,现有评估多局限于精心筛选的数据集,无法反映真实代谢组学中数据异质性和跨域分布偏移。同时,当前基准缺乏难度感知诊断,也无法揭示模型在特定计算或数据约束下的表现。为此,我们提出 FlexMS,一个模块化的公开数据基准框架,通过统一公共资源中的分子编码、元数据条件、预测头和下游检索流程,实现标准的 MS/MS 预测评估。FlexMS 不仅降低新预测工具集成门槛,更在平均性能基础上引入难度感知诊断,为不同算力、数据规模及下游检索目标提供可行动的模型选择指导。最终,该框架为社区提供可复现的标准,以识别稳定算法结论与实际可行的操作点。
原文摘要 · Abstract (English)
Tandem mass spectrometry (MS/MS) is central to small molecule identification, but current deep learning systems for spectrum prediction still remain difficult to evaluate and deploy in practice. While novel architectures constantly claim state-of-the-art performance, inconsistent metadata conditioning and entangled preprocessing pipelines hinder fair architectural comparisons. Besides, existing evaluations are often restricted to curated datasets, failing to capture the heterogeneity and cross-domain shifts of real-world metabolomics. Furthermore, current benchmarks lack difficulty-aware diagnostics and leave blind to how models behave under specific compute or data constraints. To address this, we present FlexMS, a modular public-data benchmark framework that standardizes MS/MS prediction across public resources while keeping molecular encoders, metadata conditioning, predictor heads, and downstream retrieval under one protocol. FlexMS establishes a fair evaluation playground which significantly lowers the barrier for integrating new predictive tools. Rather than solely optimizing for average scores, FlexMS augments aggregate accuracy with difficulty-aware diagnostics, providing actionable guidance on model selection across different compute constraints, data scales, and downstream retrieval objectives. Ultimately, FlexMS provides the community with a reproducible standard to identify which algorithmic conclusions are stable and which operating points are most viable in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。