SAFE框架提升分子生成效率,尤其擅长骨架修饰与连接子设计。
SAFE setup for generative molecular design
- 采用片段嵌入的序列附加机制,优化约束分子生成任务。
- 使用LLaMA架构+旋转位置编码,在大规模数据上表现最优。
- 相比SMILES模型,对骨架修饰和连接子设计更有效,尤其配合BRICS分解。
基于SMILES的分子生成模型在药物设计中至关重要,但在片段约束任务中面临挑战。为此,最近提出了基于序列附加的片段嵌入(SAFE)表示方法,以简化此类任务。本研究探讨了训练SAFE生成模型的最佳配置,重点关注数据集规模、随机化数据增强、模型架构及键断裂算法的影响。结果表明,更大、更多样化的数据集能显著提升性能,采用旋转位置编码的LLaMA架构表现最为稳健。基于SAFE的模型在骨架修饰和连接子设计任务中持续优于SMILES基方法,尤其在使用BRICS分解时效果最佳。这些发现揭示了影响SAFE生成模型效能的关键因素。
原文摘要 · Abstract (English)
SMILES-based molecular generative models have been pivotal in drug design but face challenges in fragment-constrained tasks. To address this, the Sequential Attachment-based Fragment Embedding (SAFE) representation was recently introduced as an alternative that streamlines those tasks. In this study, we investigate the optimal setups for training SAFE generative models, focusing on dataset size, data augmentation through randomization, model architecture, and bond disconnection algorithms. We found that larger, more diverse datasets improve performance, with the LLaMA architecture using Rotary Positional Embedding proving most robust. SAFE-based models also consistently outperform SMILES-based approaches in scaffold decoration and linker design, particularly with BRICS decomposition yielding the best results. These insights highlight key factors that significantly impact the efficacy of SAFE-based generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。