统一单分子与混合物的嗅觉嵌入,提升气味表征泛化能力。
AROMMA: Unifying Olfactory Embeddings for Single Molecules and Mixtures
- 用化学基础模型和注意力聚合器统一编码单分子与双分子混合物。
- 在单分子与混合物数据集上均达最优,最高提升19.1% AUROC。
- 适合从事气味生成、分子设计与跨模态表征的研究者。
公开的嗅觉数据集规模小且分散于单分子与混合物之间,限制了通用气味表征的学习。现有方法或仅学习单分子嵌入,或通过相似性/成对标签预测处理混合物,导致表征分离且未对齐。本文提出AROMMA框架,学习单分子与双分子混合物的统一嵌入空间。每个分子由化学基础模型编码,混合物通过基于注意力的聚合器构建,保证置换不变性与不对称分子相互作用。进一步通过知识蒸馏与类别感知伪标签对齐气味描述集,补充缺失的混合物标注。AROMMA在单分子与分子对数据集上均取得最佳性能,最多提升19.1% AUROC,证明了在两个领域中的鲁棒泛化能力。
原文摘要 · Abstract (English)
Public olfaction datasets are small and fragmented across single molecules and mixtures, limiting learning of generalizable odor representations. Recent works either learn single-molecule embeddings or address mixtures via similarity or pairwise label prediction, leaving representations separate and unaligned. In this work, we propose AROMMA, a framework that learns a unified embedding space for single molecules and two-molecule mixtures. Each molecule is encoded by a chemical foundation model and the mixtures are composed by an attention-based aggregator, ensuring both permutation invariance and asymmetric molecular interactions. We further align odor descriptor sets using knowledge distillation and class-aware pseudo-labeling to enrich missing mixture annotations. AROMMA achieves state-of-the-art performance in both single-molecule and molecule-pair datasets, with up to 19.1% AUROC improvement, demonstrating a robust generalization in two domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。