用专用于分子的训练框架,让小模型在药物发现上超越大模型。
MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery
- 构建分子专用训练框架,教模型理解分子语言
- 小模型在5类药研任务中逼近专家级表现且更高效
- 适合需要高效精准分子建模的研究者
通用大语言模型依赖上下文学习,在药物发现任务中难以保证科学理解与性能。单纯增大模型规模或引入推理标记,并不能带来显著提升。为此,我们提出MMAI Gym for Science,一个集分子数据格式、模态与任务特定推理、训练及评估方案于一体的综合性平台,旨在教会基础模型掌握‘分子语言’以解决实际药物发现问题。利用该平台,我们训练出一种高效的Liquid Foundation Model(LFM),结果表明,经过专门训练的小型基础模型可在分子基准测试中显著优于更大规模的通用或专用模型。在分子优化、ADMET性质预测、逆合成、药物-靶点活性预测及官能团推理等关键药物发现任务中,该模型达到接近专业级表现,在多数场景下超越大型模型,同时保持更高效率和更广适用性。
原文摘要 · Abstract (English)
General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and performance required for drug discovery tasks. Simply increasing model size or introducing reasoning tokens does not yield significant performance gains. To address this gap, we introduce the MMAI Gym for Science, a one-stop shop molecular data formats and modalities as well as task-specific reasoning, training, and benchmarking recipes designed to teach foundation models the 'language of molecules' in order to solve practical drug discovery problems. We use MMAI Gym to train an efficient Liquid Foundation Model (LFM) for these applications, demonstrating that smaller, purpose-trained foundation models can outperform substantially larger general-purpose or specialist models on molecular benchmarks. Across essential drug discovery tasks - including molecular optimization, ADMET property prediction, retrosynthesis, drug-target activity prediction, and functional group reasoning - the resulting model achieves near specialist-level performance and, in the majority of settings, surpasses larger models, while remaining more efficient and broadly applicable in the domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。