arXiv:2502.13449cs.LGphysics.chem-ph2025-02NeurIPS被引 23

Mol-LLaMA让分子语言模型具备通用理解与推理能力,助力药物研发。

Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model

  • 融合多源分子编码信息,提升特征捕捉能力
  • 在分子属性预测任务中表现优于基线模型
  • 适合化学、药学研究者用于分子分析辅助

分子理解是揭示生物机制和推动药物发现的关键,需跨化学与生物学的综合知识。尽管大型分子语言模型在任务迁移上已取得显著进展,但受限于知识与推理能力,常难以准确分析分子特性。为此,我们提出Mol-LLaMA,一种聚焦分子核心知识、具备可解释性与推理能力的大规模分子语言模型。我们设计了涵盖基本分子特征的关键数据类型,并构建模块以整合不同分子编码器的互补信息,充分发挥各类表征优势。实验结果表明,Mol-LLaMA能够有效理解分子的通用特征并生成有意义响应,展现出作为通用分子分析助手的潜力。项目页面:https://mol-llama.github.io/

原文摘要 · Abstract (English)

Understanding molecules is key to understanding organisms and driving advances in drug discovery, requiring interdisciplinary knowledge across chemistry and biology. Although large molecular language models have achieved notable success in task transfer, they often struggle to accurately analyze molecular features due to limited knowledge and reasoning capabilities. To address this issue, we present Mol-LLaMA, a large molecular language model that grasps the general knowledge centered on molecules and exhibits explainability and reasoning ability. To this end, we design key data types that encompass the fundamental molecular features, taking into account the essential abilities for molecular reasoning. Further, to improve molecular understanding, we propose a module that integrates complementary information from different molecular encoders, leveraging the distinct advantages of molecular representations. Our experimental results demonstrate that Mol-LLaMA is capable of comprehending the general features of molecules and providing informative responses, implying its potential as a general-purpose assistant for molecular analysis. Our project page is at https://mol-llama.github.io/.

分子语言模型药物发现大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。