通过多层次特征映射提升分子嗅觉预测精度
Multi-Hierarchical Fine-Grained Feature Mapping Driven by Feature Contribution for Molecular Odor Prediction
- 设计原子级细粒度特征提取模块,捕捉关键嗅觉特征
- 引入动态重要性学习机制,显著提升模型区分能力
- 结合化学先验损失函数,有效缓解类别不平衡问题
分子嗅觉预测是利用分子结构预测其气味的过程。尽管准确预测仍具挑战性,但人工智能模型可提供潜在气味建议。现有方法多依赖基础描述符或手工指纹,表达能力有限且易受类别严重不平衡影响,制约模型训练效果。为此,本文提出一种由特征贡献驱动的分层多特征映射网络(HMFNet)。具体地,设计局部多层次特征提取模块(LMFE),在原子层面进行深度特征提取,捕获对嗅觉预测至关重要的细节特征;为增强判别性原子特征提取,集成谐波调制特征映射(HMFM)模块,动态学习特征重要性与频率调制,提升模型对相关模式的捕捉能力;此外,设计全局多层次特征提取模块(GMFE),从分子图拓扑中学习全局特征,充分挖掘全局信息以增强模型判别力。为进一步缓解类别不平衡问题,提出化学先验损失(CIL)。实验结果表明,该方法在多种深度学习模型上均显著提升性能,展现出推动分子结构表征发展和加速人工智能驱动技术进步的潜力。
原文摘要 · Abstract (English)
Molecular odor prediction is the process of using a molecule's structure to predict its smell. While accurate prediction remains challenging, AI models can suggest potential odors. Existing methods, however, often rely on basic descriptors or handcrafted fingerprints, which lack expressive power and hinder effective learning. Furthermore, these methods suffer from severe class imbalance, limiting the training effectiveness of AI models. To address these challenges, we propose a Feature Contribution-driven Hierarchical Multi-Feature Mapping Network (HMFNet). Specifically, we introduce a fine-grained, Local Multi-Hierarchy Feature Extraction module (LMFE) that performs deep feature extraction at the atomic level, capturing detailed features crucial for odor prediction. To enhance the extraction of discriminative atomic features, we integrate a Harmonic Modulated Feature Mapping (HMFM). This module dynamically learns feature importance and frequency modulation, improving the model's capability to capture relevant patterns. Additionally, a Global Multi-Hierarchy Feature Extraction module (GMFE) is designed to learn global features from the molecular graph topology, enabling the model to fully leverage global information and enhance its discriminative power for odor prediction. To further mitigate the issue of class imbalance, we propose a Chemically-Informed Loss (CIL). Experimental results demonstrate that our approach significantly improves performance across various deep learning models, highlighting its potential to advance molecular structure representation and accelerate the development of AI-driven technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。