用少量数据生成可合成的全新香味分子,突破传统设计瓶颈。
QSAR-Guided Generative Framework for the Discovery of Synthetically Viable Odorants
- VAE结合外部QSAR模型,在小样本下学习气味特性
- 生成结构100%有效,94.8%为新结构,FCD仅6.96优于基线
- 74.4%候选物具有全新骨架,适合创新香料研发
新香味分子的发现对香精香料行业至关重要,但高效探索庞大化学空间以识别具有理想嗅觉特性的结构仍是重大挑战。生成式人工智能为从头分子设计提供了前景,但通常需要大量分子数据进行训练。为此,我们提出一种结合变分自编码器(VAE)与定量构效关系(QSAR)模型的框架,可在有限训练集下生成新型气味分子。VAE通过自监督学习从ChemBL数据库中习得SMILES语法规则,其训练目标还引入了来自外部QSAR模型的损失项,以根据气味概率结构化潜在表示。尽管VAE在内部一致性上表现良好,但在外部未见的真实数据集(Unique Good Scents)上的验证表明,模型生成的结构语法正确率高达100%(通过拒绝采样实现),且94.8%为唯一结构。潜在空间有效按气味可能性组织,生成分子与已知气味分子间的弗雷切特化学网络距离(FCD)约为6.96,远优于ChemBL基线的约21.6。通过Bemis-Murcko骨架分析显示,74.4%的候选物拥有不同于训练数据的新核心骨架,表明模型在化学空间中进行了广泛探索,而非简单衍化已有气味分子。
原文摘要 · Abstract (English)
The discovery of novel odorant molecules is key for the fragrance and flavor industries, yet efficiently navigating the vast chemical space to identify structures with desirable olfactory properties remains a significant challenge. Generative artificial intelligence offers a promising approach for \textit{de novo} molecular design but typically requires large sets of molecules to learn from. To address this problem, we present a framework combining a variational autoencoder (VAE) with a quantitative structure-activity relationship (QSAR) model to generate novel odorants from limited training sets of odor molecules. The self-supervised learning capabilities of the VAE allow it to learn SMILES grammar from ChemBL database, while its training objective is augmented with a loss term derived from an external QSAR model to structure the latent representation according to odor probability. While the VAE demonstrated high internal consistency in learning the QSAR supervision signal, validation against an external, unseen ground truth dataset (Unique Good Scents) confirms the model generates syntactically valid structures (100\% validity achieved via rejection sampling) and 94.8\% unique structures. The latent space is effectively structured by odor likelihood, evidenced by a Fréchet ChemNet Distance (FCD) of $\approx$ 6.96 between generated molecules and known odorants, compared to $\approx$ 21.6 for the ChemBL baseline. Structural analysis via Bemis-Murcko scaffolds reveals that 74.4\% of candidates possess novel core frameworks distinct from the training data, indicating the model performs extensive chemical space exploration beyond simple derivatization of known odorants. Generated candidates display physicochemical properties ....
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。