arXiv:2506.15792cs.LGphysics.chem-ph2025-06被引 14

用经典分子描述符训练大模型,预测性能超越传统方法。

Deep Learning Foundation Models from Classical Molecular Descriptors

  • 基于低噪声分子描述符构建大规模预训练模型
  • 在58个数据集上胜率达75%至97%,超越随机森林等基线
  • 适合需要高精度分子性质预测的研究者使用

快速准确的数据驱动分子性质预测对众多化学领域的发展至关重要。尽管深度学习近年备受关注,但在真实世界、小样本场景下仍难以超越经典机器学习方法。本研究提出CheMeleon,一个约1000万参数的奠基模型,使定向消息传递神经网络首次超越经典方法。在Polaris和MoleculeACE的58个基准数据集上评估,其在Polaris任务中胜率达75%,优于随机森林(68%)、fastprop(36%)和Chemprop(32%);在MoleculeACE测试中胜率高达97%,超过随机森林(50%)及其他基础模型。与依赖嘈杂实验数据或有偏量子模拟的传统预训练不同,CheMeleon利用低噪声分子描述符学习丰富且高度可迁移的分子表征,为奠基模型预训练提供新路径。

原文摘要 · Abstract (English)

Fast and accurate data-driven prediction of molecular properties is pivotal to scientific advancements across myriad chemical domains. Deep learning methods have recently garnered much attention, despite their inability to outperform classical machine learning methods when tested on practical, real-world benchmarks with limited training data. This study seeks to bridge this gap with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods. Evaluated on 58 benchmark datasets from Polaris and MoleculeACE, CheMeleon achieves a win rate of 75% on Polaris tasks, outperforming baselines like Random Forest (68%), fastprop (36%), and Chemprop (32%), and a 97% win rate on MoleculeACE assays, surpassing Random Forest (50%) and other foundation models. Unlike conventional pre-training approaches that rely on noisy experimental data or biased quantum mechanical simulations, CheMeleon utilizes low-noise molecular descriptors to learn rich and highly transferable molecular representations, suggesting a new avenue for foundation model pre-training.

分子建模深度学习基础模型特征表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。