arXiv:2604.26498cs.LGq-bio.QM2026-04被引 2

大模型未必更优,小模型在药物发现中表现更稳定。

Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

论文配图:Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
图 1 · 摘自论文原文
  • 对比156项任务,经典机器学习仍胜出近半数。
  • 图神经网络和序列模型仅在复杂场景下竞争力强。
  • 模型效果取决于任务与验证方式,而非单纯看大小。

分子基础模型和大语言模型的快速发展推动了药物发现领域对模型规模的过度关注,认为更大模型必胜。我们通过26个ADME、毒性和生物活性预测任务,覆盖165,541个化合物标签记录,评估了不同模型在随机、Murcko骨架和结构分离三种5折交叉验证下的表现。结果表明:经典机器学习(ML)在156项任务比较中以47.4%的占比领先,其次是预训练分子序列模型(28.8%)、图神经网络(21.8%)和基于LLM的SAR基线(1.9%)。经典ML在随机划分下表现最佳,整体仍是最大赢家。图神经网络和序列模型在更难的划分场景下有竞争力,但在固定最终窗口读出时胜率下降,显示对训练设置敏感。配对自举分析表明,模型间微小差异不应视为决定性优势。训练集中的结构-活性关系知识虽提升GPT5.5-SAR和Opus4.7-SAR性能,但无法替代监督学习模型。专用小型模型依然高效,预测能力取决于模型、任务与验证场景的匹配度,而非模型规模。

原文摘要 · Abstract (English)

The rapid growth of molecular foundation models and large language models (LLMs) has encouraged a scale centred view of AI in drug discovery, in which larger pretrained models are expected to supersede compact cheminformatics models. We test this assumption across 26 ADME, toxicity and bioactivity endpoints, covering 165,541 endpoint level compound label records. The benchmark contains 78 endpoint and split entries evaluated under random, Murcko scaffold and structure separated 5-fold cross validation protocols, representing increasing chemical generalization difficulty. Across 156 task and metric comparisons, classical machine learning (ML) provides the largest share of best performing entries (47.4%), followed by pretrained molecular sequence models (28.8%), graph neural networks (21.8%) and LLM based SAR baselines (1.9%). Classical ML dominates random split interpolation and remains the largest winner family overall. GNN and sequence models are competitive in selected harder splits, but their strict winner shares decrease under a fixed final-window readout, indicating sensitivity to training settings and model selection. Paired bootstrap analyses show that small numerical differences between individual models should not be read as decisive victories. SAR knowledge from training folds improves GPT5.5-SAR and Opus4.7-SAR metrics but does not make rule based reasoning a universal substitute for supervised predictors. Compact specialized models remain highly effective, and predictive performance depends on the fit among model, task and validation scenario, not on scale alone.

药物发现模型规模机器学习分子预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。