构建多层级分子知识数据集,提升化学大模型理解与生成能力
KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
- 构建10万条多层级分子标注数据,连接分子与文本描述
- 提出化学信息丰富的分子表示方法,改进现有表征策略
- 在分子理解与生成任务中表现领先,适合药物研发等场景
分子大语言模型因其在分子应用中的巨大潜力而受到广泛关注。然而,当前模型在分子理解方面仍存在显著局限,主要源于预训练阶段文本描述不足和分子表征策略不佳。为此,我们提出了KnowMol-100K,一个包含10万条跨多层级的精细分子标注的大型数据集,弥合了分子与文本描述之间的差距。同时,我们提出了一种化学信息丰富的分子表示方法,有效解决了现有表征策略的不足。基于这些创新,我们开发了当前最先进的多模态分子大语言模型KnowMol。大量实验表明,KnowMol在分子理解与生成任务中均取得卓越性能。
原文摘要 · Abstract (English)
The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understanding molecules due to inadequate textual descriptions and suboptimal molecular representation strategies during pretraining. To address these challenges, we introduce KnowMol-100K, a large-scale dataset with 100K fine-grained molecular annotations across multiple levels, bridging the gap between molecules and textual descriptions. Additionally, we propose chemically-informative molecular representation, effectively addressing limitations in existing molecular representation strategies. Building upon these innovations, we develop KnowMol, a state-of-the-art multi-modal molecular large language model. Extensive experiments demonstrate that KnowMol achieves superior performance across molecular understanding and generation tasks. GitHub: https://github.com/yzf-code/KnowMol Huggingface: https://hf.co/datasets/yzf1102/KnowMol-100K
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。