大模型在化学领域预训练效果未必能转化为下游任务提升。
How Well Do Large-Scale Chemical Language Models Transfer to Downstream Tasks?
- 系统放大模型、数据和算力,测试预训练与下游性能关系。
- 预训练损失持续下降,但下游任务表现几乎停滞甚至下降。
- 现有评估指标无法预测实际应用效果,需关注任务特性。
基于大规模分子数据预训练的化学语言模型(CLMs)广泛用于分子性质预测。然而,普遍认为增大模型规模、数据量和计算资源可同时降低预训练损失并提升下游性能这一观点,在化学领域尚未得到系统验证。本文通过系统性地扩大训练资源,预训练多个CLM,并在多样化的分子性质预测(MPP)任务上评估其迁移能力。结果发现:尽管预训练损失随资源增加持续下降,下游任务性能却提升有限;基于海森矩阵或损失景观的替代评估指标也难以准确预测下游表现。我们进一步识别出下游性能饱和或退化的条件,并通过参数空间可视化分析了不同任务下的失效模式。研究揭示了预训练评估与实际应用之间的差距,强调应建立考虑下游任务特性的模型选择与评估策略。
原文摘要 · Abstract (English)
Chemical Language Models (CLMs) pre-trained on large scale molecular data are widely used for molecular property prediction. However, the common belief that increasing training resources such as model size, dataset size, and training compute improves both pretraining loss and downstream task performance has not been systematically validated in the chemical domain. In this work, we evaluate this assumption by pretraining CLMs while scaling training resources and measuring transfer performance across diverse molecular property prediction (MPP) tasks. We find that while pretraining loss consistently decreases with increased training resources, downstream task performance shows limited improvement. Moreover, alternative metrics based on the Hessian or loss landscape also fail to estimate downstream performance in CLMs. We further identify conditions under which downstream performance saturates or degrades despite continued improvements in pretraining metrics, and analyze the underlying task dependent failure modes through parameter space visualizations. These results expose a gap between pretraining based evaluation and downstream performance, and emphasize the need for model selection and evaluation strategies that explicitly account for downstream task characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。