arXiv:2410.11914q-bio.QMcs.LG2024-10中稿 · as a short paper a…被引 1

用更大知识图谱提升分子属性预测效果

Large-Scale Knowledge Integration for Enhanced Molecular Property Prediction

  • 引入2840个功能基团的ChEBI知识图谱增强模型
  • 在14个数据集中有9个表现优于原模型
  • 适合药物发现与材料科学中的分子表征研究

在分子属性上预训练机器学习模型已被证明能生成稳健且可泛化的表征,对药物发现和材料科学至关重要。尽管现有工作多聚焦于数据驱动方法,KANO模型提出了一种新范式,通过知识增强预训练实现突破。本文在原有基础上,集成包含2,840个功能基团的大型ChEBI知识图谱,远超原版使用的82个。我们探索了“替换”与“融合”两种方式将该知识融入KANO框架。结果表明,在14个分子属性预测数据集中有9个表现更优,验证了使用更大、更多样化的功能基团集合对提升分子表征能力的关键作用。代码已开源:github.com/Yasir-Ghunaim/KANO-ChEBI。

原文摘要 · Abstract (English)

Pre-training machine learning models on molecular properties has proven effective for generating robust and generalizable representations, which is critical for advancements in drug discovery and materials science. While recent work has primarily focused on data-driven approaches, the KANO model introduces a novel paradigm by incorporating knowledge-enhanced pre-training. In this work, we expand upon KANO by integrating the large-scale ChEBI knowledge graph, which includes 2,840 functional groups -- significantly more than the original 82 used in KANO. We explore two approaches, Replace and Integrate, to incorporate this extensive knowledge into the KANO framework. Our results demonstrate that including ChEBI leads to improved performance on 9 out of 14 molecular property prediction datasets. This highlights the importance of utilizing a larger and more diverse set of functional groups to enhance molecular representations for property predictions. Code: github.com/Yasir-Ghunaim/KANO-ChEBI

分子表征知识图谱属性预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。