arXiv:2606.19624cs.LG2026-06被引 3

揭露质谱分子发现模型评估中的三大漏洞,推动更可信的基准测试。

MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery

论文配图:MassSpecGym in the Wild: Uncovering and Correcting Evaluation Pitfalls in AI-Driven Molecule Discovery
图 1 · 摘自论文原文
  • 通过复现26篇论文,发现17篇存在数据泄露、捷径学习等问题
  • 验证这些错误使评估结果失真,破坏基准可靠性
  • 发布v1.5版本修复缺陷,适合质谱机器学习研究者使用

可靠的基准测试对基于串联质谱(MS/MS)的分子发现机器学习模型开发至关重要。实验设计与模型评估中的细微问题会降低基准的可信度,导致错误结论。我们对近期MS/MS机器学习文献中的评估问题进行了全面审查,以标准的MassSpecGym基准套件为案例,揭示其影响。在采用该基准首年报告结果的26篇论文中,至少有17篇存在评估缺陷。我们识别出三类失败:(i) 数据泄露,(ii) 捷径学习,(iii) 实现错误与指标偏差。通过大量实验与代码复现,量化了这些问题的影响,证明它们破坏了MassSpecGym原本旨在建立的评估标准。我们提炼出可推广至各类MS/MS挑战、基准及自定义评估设置的建议。同时发布MassSpecGym v1.5,整合上述建议,修复已识别的问题模式。v1.5可在https://github.com/pluskal-lab/MassSpecGym公开获取。

原文摘要 · Abstract (English)

Reliable benchmarking is critical for developing machine learning models for tandem mass spectrometry (MS/MS) based molecule discovery. Subtle issues in experimental design and model evaluation procedures can degrade the trustworthiness of such benchmarks and lead to erroneous conclusions. We conduct a thorough review of model evaluation issues in the recent MS/MS machine learning literature, using the standard MassSpecGym benchmark suite as a case study to illustrate the impact of these issues. We find evaluation issues in at least 17 of 26 papers reporting MassSpecGym benchmark results in the first year of its adoption. We isolate three classes of failures: (i) data leakage, (ii) shortcut learning, and (iii) implementation bugs and metric divergence. Through extensive experimentation and code replication, we quantify the impact of these issues and show how they corrupt the evaluation standards MassSpecGym was designed to enforce. We distill our findings into recommendations generalizable to MS/MS challenges, benchmarks, and custom evaluation setups. We also release MassSpecGym v1.5, an implementation of our recommendations in the MassSpecGym benchmarking suite which addresses the failure modes identified in this audit. MassSpecGym v1.5 is publicly available at https://github.com/pluskal-lab/MassSpecGym.

质谱分析模型评估基准测试机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。