arXiv:2601.22757cs.LGq-bio.BM2026-01被引 1

揭示分子语言模型的缩放规律,明确模型、数据与表示的影响。

Unveiling Scaling Behaviors in Molecular Language Models: Effects of Model Size, Data, and Representation

  • 系统控制算力,分别测试模型大小、数据量和分子表示的影响。
  • 发现预训练与下游任务均有清晰缩放规律,表现随数据量提升显著。
  • 揭示不同分子表示导致以往研究结果不一致,适合模型优化研究者。

分子生成模型通常在分子字符串表示上采用类似GPT的语言建模方法,在大规模数据集和模型规模下展现出良好性能。然而,在固定计算预算下,这些模型是否遵循可预测的缩放规律仍不明确且存在争议,这对合理分配模型规模、数据量和分子表示资源至关重要。本研究系统地探究了分子语言模型在预训练和下游任务中的缩放行为。我们训练了300个模型,执行超过10,000次实验,在严格控制计算预算的前提下,独立变化模型规模、训练词元数量和分子表示方式。结果表明,分子模型在预训练和下游迁移任务中均存在清晰的缩放规律,分子表示对性能影响显著,并解释了此前观察到的分子生成缩放行为不一致现象。此外,我们公开发布了迄今为止最大的分子语言模型库,以促进未来研究。代码与模型已发布于 https://github.com/SZU-ADDG/MLM-Scaling。

原文摘要 · Abstract (English)

Molecular generative models, often employing GPT-style language modeling on molecular string representations, have shown promising capabilities when scaled to large datasets and model sizes. However, it remains unclear and subject to debate whether these models adhere to predictable scaling laws under fixed computational budgets, which is a crucial understanding for optimally allocating resources between model size, data volume, and molecular representation. In this study, we systematically investigate the scaling behavior of molecular language models across both pretraining and downstream tasks. We train 300 models and conduct over 10,000 experiments, rigorously controlling compute budgets while independently varying model size, number of training tokens, and molecular representation. Our results demonstrate clear scaling laws in molecular models for both pretraining and downstream transfer, reveal the substantial impact of molecular representation on performance, and explain previously observed inconsistencies in scaling behavior for molecular generation. Additionally, we publicly release the largest library of molecular language models to date to facilitate future research and development. Code and models are available at https://github.com/SZU-ADDG/MLM-Scaling.

分子生成缩放规律语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。