探究化学语言模型如何学习分子结构,发现预训练提升结构感知能力。
Probing Chemical Language Models: Effects of Pre-training and Fine-tuning

- 通过探测78种分子子结构,分析模型各层对化学信息的编码
- 预训练显著增强模型对分子结构的感知,尤其在高层特征中
- 微调会优先改变与任务相关的子结构,符合化学理论
化学语言模型(CLMs)通常使用线性表示如SMILES进行训练,但其编码的化学有意义子结构仍不明确。为深入理解CLMs,我们系统性地探测了8个预训练模型和6个随机初始化模型中的78种分子子结构。此外,研究了在化学下游任务上微调对分子子结构表征的影响。结果表明,预训练总体提升了CLMs对分子结构的感知能力,尤其在模型高层。随机初始化模型在第一层已能较好编码环状结构。对两个化学下游任务的分析进一步显示,微调更显著影响与任务相关的分子子结构,表明表征变化遵循化学理论。
原文摘要 · Abstract (English)
Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode. To foster a better understanding of CLMs, we conduct a systematic study and probe for 78 molecular substructures across eight pre-trained and six randomly initialized models. We furthermore study how fine-tuning on chemical downstream tasks affects the learned representations of molecular substructures. Our results show that pre-training generally improves molecular structure awareness of CLMs, particularly in the upper layers. Moreover, randomly initialized models already encode ring structures well in the first layer. Our analysis on two chemical downstream tasks further reveals that, interestingly, fine-tuning affects task-relevant molecular substructures more than others, indicating that the changes in the representations follow chemical theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。