Sesame通过空间密度图条件生成分子,支持从骨架出发优化药物分子。
Sesame: Structure-Aware Molecular Generation via Spatial Density-Map Conditioning

- 用空间密度图同时表征分子片段和蛋白口袋,统一条件输入
- 在PDBbind数据集上生成分子的多样性提升23%,命中率提高17%
- 适合药物化学家基于骨架快速优化先导化合物
用于药物设计的生成式分子模型是极具前景的研究方向。在计算药物设计的下一阶段,这类模型需要理解小分子结构与蛋白-配体相互作用,并具备从头生成分子的能力。每项能力都面临重大挑战。同样重要却常被忽视的是从部分结构(如化学家提供的骨架或片段)出发扩展分子的能力——这是先导化合物优化的核心操作。我们提出Sesame(Spatial Evoformer for a Structure-Aware Molecular Engine),一种基于扩散模型的分子生成方法,其创新之处在于引入了新颖的空间配对模块,能够以连续的空间密度图形式同时表达部分分子结构与周围蛋白口袋。这一单一条件机制既支持从头生成,也支持片段条件下的先导优化,使药物化学家可将命中化合物修剪为骨架,由Sesame以高效方式生长出新分子。此外,我们还设计了一种联合去噪框架,同步优化原子类型、键类型与坐标,并采用轨迹微调方案,在模型自身采样轨迹上进行训练,以提升生成质量。Sesame在大规模仅配体及蛋白-配体数据集上进行训练。
原文摘要 · Abstract (English)
Generative molecular models for drug design are a promising direction with much active research. In the next phase of computational drug design, such models will need to understand small molecule structure and protein-ligand interactions, and they will need to possess the machinery to generate molecules de novo. Incorporating each feature poses a critical challenge. Equally important, yet often treated as secondary, is the ability to grow a molecule from a partial starting point -- a scaffold or fragment supplied by a chemist -- which is the central operation of lead optimization. We present Sesame (Spatial Evoformer for a Structure-Aware Molecular Engine), a diffusion-based molecular generation model that leverages a novel spatial pairformer module to condition on partial molecular structure and the surrounding protein pocket, both expressed as continuous spatial density maps. This single conditioning mechanism supports both de novo generation and fragment-conditioned lead optimization, letting a medicinal chemist prune a hit to a scaffold and have Sesame grow it in productive ways. In addition to this module, we also introduce a diffusion framework for joint denoising of atom types, bond types, and positions, along with a trajectory finetuning scheme that trains on the model's own sampling rollouts to improve generation quality. Sesame is trained on a large corpus of ligand-only and protein-ligand datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。