用混合分子建模嗅觉相似性,提升预测精度
From Molecules to Mixtures: Learning Representations of Olfactory Mixture Similarity using Inductive Biases
- 分层构建:图网络+注意力+余弦预测,融合分子对气味的贡献
- 在多数据集上达到最优预测性能,泛化至未见分子与混合规模
- 适合对嗅觉感知、化学表征感兴趣的科研人员
嗅觉——分子如何被人类感知为气味——仍不明确。最近提出的主嗅觉地图(POM)实现了单个化合物嗅觉特性的数字化。然而,真实气味并非纯分子,而是复杂分子混合物,其表征仍相对未被充分探索。本文提出POMMix,扩展POM以表示混合物。该表征基于问题空间对称性进行分层构建:(1) 使用图神经网络生成分子嵌入;(2) 通过注意力机制将分子表示聚合为混合物表示;(3) 采用余弦预测头编码混合物嵌入空间中的嗅觉感知距离。POMMix在多个数据集上实现最先进的预测性能,并评估了在未见分子和不同混合物规模下的泛化能力。本工作推动了嗅觉数字化进程,凸显了领域知识与深度学习在低数据场景下构建表达性强表征的协同效应。
原文摘要 · Abstract (English)
Olfaction -- how molecules are perceived as odors to humans -- remains poorly understood. Recently, the principal odor map (POM) was introduced to digitize the olfactory properties of single compounds. However, smells in real life are not pure single molecules, but complex mixtures of molecules, whose representations remain relatively under-explored. In this work, we introduce POMMix, an extension of the POM to represent mixtures. Our representation builds upon the symmetries of the problem space in a hierarchical manner: (1) graph neural networks for building molecular embeddings, (2) attention mechanisms for aggregating molecular representations into mixture representations, and (3) cosine prediction heads to encode olfactory perceptual distance in the mixture embedding space. POMMix achieves state-of-the-art predictive performance across multiple datasets. We also evaluate the generalizability of the representation on multiple splits when applied to unseen molecules and mixture sizes. Our work advances the effort to digitize olfaction, and highlights the synergy of domain expertise and deep learning in crafting expressive representations in low-data regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。