arXiv:2410.09240cs.LGcs.CL2024-10AAAI被引 6

用点云编码器让语言模型读懂分子三维结构,支持多任务生成。

nach0-pc: Multi-task Language Model with Molecular Point Cloud Encoder

  • 用分子点云编码器实现无序、紧凑的三维结构表示。
  • 在多个生成任务中性能媲美扩散模型,且训练推理更快。
  • 适合需要处理分子3D结构的药物设计研究者使用。

近期进展将语言模型(LM)引入药物发现流程,但现有模型多依赖SMILES和SELFIES等化学字符串表示,缺乏对药物发现至关重要的空间信息。将化学3D结构转为文本格式又存在长度过长、原子连接信息不足等问题。为此,我们提出nach0-pc,结合领域专用编码器与文本表示,有效处理原子空间排列。该方法采用分子点云编码器实现简洁、顺序无关的结构表示,并设计新型预训练方案,从空间分子结构数据集中蒸馏知识。在单任务与多任务框架下微调后,nach0-pc在多个经典的空间分子生成任务中生成样本质量可与扩散模型比肩。值得注意的是,本模型为多任务设计,突破了扩散模型仅限单任务的局限;同时具备处理点云数据的能力,克服了语言模型因内存限制无法直接处理点云的问题。这使得模型训练与推理时间显著降低,性能保持相当。

原文摘要 · Abstract (English)

Recent advancements have integrated Language Models (LMs) into a drug discovery pipeline. However, existing models mostly work with SMILES and SELFIES chemical string representations, which lack spatial features vital for drug discovery. Additionally, attempts to translate chemical 3D structures into text format encounter issues such as excessive length and insufficient atom connectivity information. To address these issues, we introduce nach0-pc, a model combining domain-specific encoder and textual representation to handle spatial arrangement of atoms effectively. Our approach utilizes a molecular point cloud encoder for concise and order-invariant structure representation. We introduce a novel pre-training scheme for molecular point clouds to distillate the knowledge from spatial molecular structures datasets. After fine-tuning within both single-task and multi-task frameworks, nach0-pc demonstrates performance comparable with other diffusion models in terms of generated samples quality across several established spatial molecular generation tasks. Notably, our model is a multi-task approach, in contrast to diffusion models being limited to single tasks. Additionally, it is capable of processing point cloud-related data, which language models are not capable of handling due to memory limitations. These lead to our model having reduced training and inference time while maintaining on par performance.

分子生成点云编码多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。