系统实验揭示分子属性预测中语言模型表现受数据量、模型规模等多重因素影响。
BERTology of Molecular Property Prediction
- 通过数百次受控实验,分析数据量、模型大小和标准化对预训练效果的影响。
- 发现现有模型性能差异源于未被充分重视的训练条件差异。
- 为分子属性预测提供可复现的基准与优化方向,适合研究者参考。
化学语言模型(CLMs)已成为分子属性预测(MPP)任务中颇具前景的替代方法,但多项研究在不同基准任务上报告了不一致甚至矛盾的性能结果。本研究开展并分析了数百次精心设计的受控实验,系统考察了数据集规模、模型规模及标准化等因素对CLMs在预训练与微调阶段表现的影响。由于编码器仅结构的掩码语言模型尚无明确的扩展规律,本文旨在提供全面的数值证据,并深入理解影响CLMs在MPP任务中表现的潜在机制,其中一些在现有文献中几乎被完全忽略。
原文摘要 · Abstract (English)
Chemical language models (CLMs) have emerged as promising competitors to popular classical machine learning models for molecular property prediction (MPP) tasks. However, an increasing number of studies have reported inconsistent and contradictory results for the performance of CLMs across various MPP benchmark tasks. In this study, we conduct and analyze hundreds of meticulously controlled experiments to systematically investigate the effects of various factors, such as dataset size, model size, and standardization, on the pre-training and fine-tuning performance of CLMs for MPP. In the absence of well-established scaling laws for encoder-only masked language models, our aim is to provide comprehensive numerical evidence and a deeper understanding of the underlying mechanisms affecting the performance of CLMs for MPP tasks, some of which appear to be entirely overlooked in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。