用专家与数据驱动的气味分类法提升分子结构预测气味的准确率
Exploring Molecular Odor Taxonomies for Structure-based Odor Predictions using Machine Learning
- 构建了基于语义和数据聚类的双重气味分类体系
- 两类分类法均显著优于随机分组,提升模型预测性能
- 公开777个描述符的多层专家分类与完整数据集,助力嗅觉研究
从分子结构预测气味的核心挑战在于对气味空间理解有限及结构-气味关系复杂。本文表明,通过结合专家定义与数据驱动的气味分类体系,可显著提升机器学习模型的预测性能。专家分类基于语义与感知相似性,数据驱动分类则基于气味描述符在数据集中共现模式的聚类。两类分类均优于无关系的随机分组,且在不同气味类别上表现一致。通过跨类别预测效果评估分类质量,并开展深入误差分析,揭示气味-结构关系的复杂性,发现部分香水用气味剂存在的分类不一致性。数据驱动分类可用于检验与优化专家分类,更清晰刻画分子气味空间。本文提供包含777个描述符的多层专家分类体系及完整数据集,推动社区共建嗅觉分子基础研究。
原文摘要 · Abstract (English)
One of the key challenges to predict odor from molecular structure is unarguably our limited understanding of the odor space and the complexity of the underlying structure-odor relationships. Here, we show that the predictive performance of machine learning models for structure-based odor predictions can be improved using both, an expert and a data-driven odor taxonomy. The expert taxonomy is based on semantic and perceptual similarities, while the data-driven taxonomy is based on clustering co-occurrence patterns of odor descriptors directly from the prepared dataset. Both taxonomies improve the predictions of different machine learning models and outperform random groupings of descriptors that do not reflect existing relations between odor descriptors. We assess the quality of both taxonomies through their predictive performance across different odor classes and perform an in-depth error analysis highlighting the complexity of odor-structure relationships and identifying potential inconsistencies within the taxonomies by showcasing pear odorants used in perfumery. The data-driven taxonomy allows us to critically evaluate our expert taxonomy and better understand the molecular odor space. Both taxonomies as well as a full dataset are made available to the community, providing a stepping stone for a future community-driven exploration of the molecular basis of smell. In addition, we provide a detailed multi-layer expert taxonomy including a total of 777 different descriptors from the Pyrfume repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。