用图注意力模型预测分子气味,比传统方法更准更省参数。
GraphNOSE: A Graph Transformer in Olfaction

- 基于图变压器架构,融合分子位置与结构编码
- 对未知分子的气味预测准确率达84%(比现有最佳高3%)
- 可解释性强,能识别影响气味的关键化学片段
从分子结构预测嗅觉特性是化学生物信息学中的开放问题。尽管线性模型能关联分子特征与气味描述符,但在面对新型化学骨架、极端分子量或复杂气味混合物时往往失效。为此,我们提出GraphNOSE,一个开源的图变压器框架,可从SMILES字符串中预测单个分子及二元混合物的多标签气味描述符。通过在基于Transformer的图架构中集成位置与结构编码,GraphNOSE仅需标准图神经网络(GNN)六分之一的参数,却平均在受试者工作特征曲线下面积(AUROC)上领先线性模型、分子语言模型嵌入、分子指纹和基线GNN 4.52%(p < 0.01)。在分布外化合物(OOD)上的表现达到84%的AUROC,显著优于当前最优的嗅觉分布外模型Open-POM(81%,p < 0.001),并揭示了线性模型实际失效的条件。最后,我们运用可解释人工智能(XAI)方法识别驱动气味预测的子结构与分子特征,结果与化学直觉一致且扎根于模型学习表征。这些成果确立了GraphNOSE作为可扩展、可解释的嗅觉预测架构,能泛化至当前感知数据库中代表性不足的结构差异化合物。
原文摘要 · Abstract (English)
Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. Although linear models can link molecular features to odor descriptors, they often fail when extrapolating to novel chemical scaffolds, extreme molecular weights, or complex odor mixtures. To address this, we introduce GraphNOSE, an open-source graph transformer framework that predicts multi-label odor descriptors from simplified molecular-input line-entry system (SMILES) strings for single molecules and binary mixtures. By integrating positional and structural encodings within a transformer-based graph architecture, GraphNOSE achieves strong performance with six times fewer parameters than standard graph neural network (GNN) baseline while consistently outperforming linear models, molecular language model embeddings, molecular fingerprints, and baseline GNNs by an average area under the ROC curve (AUROC) margin of 4.52% (p < 0.01). GraphNOSE achieves an AUROC of 84% on out-of-distribution compounds (OODs). This exceeds the current state-of-the-art GNN for OOD in olfaction (Open-POM: 81%, p < 0.001), and identifies conditions under which linear models empirically fail. Finally, we apply XAI (explainable AI) methods to identify which substructures and molecular features drive odor predictions, yielding insights consistent with chemical intuition and grounded in the model's learned representations. Together, these results establish GraphNOSE as a scalable and interpretable architecture for olfactory prediction that generalizes to structurally distinct compounds underrepresented in current perceptual databases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。