用灵活方法让预训练分子图模型更好理解原子键信息。
MolGA: Molecular Graph Adaptation with Pre-trained 2D Graph Encoder
- 通过对齐策略连接拓扑与化学知识表示
- 用条件机制生成特定实例的嵌入,提升精度
- 在11个数据集上验证,适合下游分子任务
分子图表示学习广泛应用于化学与生物医学研究。尽管预训练的二维图编码器表现优异,但忽略了原子和键等子分子实例相关的丰富领域知识。虽然现有分子预训练方法将此类知识纳入预训练目标,但通常针对特定类型知识设计,难以融合多种知识。因此,复用广泛可用且经过验证的预训练2D编码器,并在下游适配中引入分子领域知识,是一种更实用的方案。本文提出MolGA,通过灵活整合多样分子领域知识,适应预训练2D图编码器至下游分子应用。首先,提出分子对齐策略,弥合预训练拓扑表示与领域知识表示之间的差距;其次,引入条件适配机制,生成实例特异性令牌,实现分子领域知识的细粒度融合;最后,在11个公开数据集上进行大量实验,验证了MolGA的有效性。
原文摘要 · Abstract (English)
Molecular graph representation learning is widely used in chemical and biomedical research. While pre-trained 2D graph encoders have demonstrated strong performance, they overlook the rich molecular domain knowledge associated with submolecular instances (atoms and bonds). While molecular pre-training approaches incorporate such knowledge into their pre-training objectives, they typically employ designs tailored to a specific type of knowledge, lacking the flexibility to integrate diverse knowledge present in molecules. Hence, reusing widely available and well-validated pre-trained 2D encoders, while incorporating molecular domain knowledge during downstream adaptation, offers a more practical alternative. In this work, we propose MolGA, which adapts pre-trained 2D graph encoders to downstream molecular applications by flexibly incorporating diverse molecular domain knowledge. First, we propose a molecular alignment strategy that bridge the gap between pre-trained topological representations with domain-knowledge representations. Second, we introduce a conditional adaptation mechanism that generates instance-specific tokens to enable fine-grained integration of molecular domain knowledge for downstream tasks. Finally, we conduct extensive experiments on eleven public datasets, demonstrating the effectiveness of MolGA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。