用合成任务提升图神经网络分子属性预测性能
All You Need Is Synthetic Task Augmentation
- 用XGBoost生成合成任务,与实验数据联合训练图变压器
- 19个分子属性任务中16个超越单任务XGBoost模型
- 无需预训练或特征注入,适合药物研发领域应用
将基于规则的模型(如随机森林)融入可微神经网络框架仍是机器学习中的开放挑战。尽管预训练模型能生成高效的分子嵌入,但通常需要大量预训练和附加技术(如后验概率)来提升性能。本研究提出一种新策略:在稀疏多任务分子性质实验目标与基于Osmordred分子描述符训练的XGBoost模型生成的合成目标上,联合训练单一图变换器神经网络。这些合成任务作为独立辅助任务。结果表明,在全部19个分子性质预测任务中均实现显著且一致的性能提升;其中16个任务的多任务图变换器优于单任务XGBoost模型。这证明合成任务增强是一种有效方法,可在无需特征注入或预训练的情况下提升神经网络在多任务分子性质预测中的表现。
原文摘要 · Abstract (English)
Injecting rule-based models like Random Forests into differentiable neural network frameworks remains an open challenge in machine learning. Recent advancements have demonstrated that pretrained models can generate efficient molecular embeddings. However, these approaches often require extensive pretraining and additional techniques, such as incorporating posterior probabilities, to boost performance. In our study, we propose a novel strategy that jointly trains a single Graph Transformer neural network on both sparse multitask molecular property experimental targets and synthetic targets derived from XGBoost models trained on Osmordred molecular descriptors. These synthetic tasks serve as independent auxiliary tasks. Our results show consistent and significant performance improvement across all 19 molecular property prediction tasks. For 16 out of 19 targets, the multitask Graph Transformer outperforms the XGBoost single-task learner. This demonstrates that synthetic task augmentation is an effective method for enhancing neural model performance in multitask molecular property prediction without the need for feature injection or pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。